Share of Voice in AI Search: How to Calculate It Correctly (and Defend the Number in a Meeting)
Learn how to calculate share of voice in AI search correctly: fix your prompt set, run repeated samples, weight by answer position, and trend it over time.

Share of Voice in AI Search: The Number Most Teams Calculate Wrong
Pew Research Center found that when a Google AI Overview appeared in March 2025, users clicked a traditional result in only 8% of visits, versus 15% without one β and clicked a link inside the summary just 1% of the time (Pew Research Center, July 2025). The answer text, not the click, now carries brand exposure. That shift is why share of voice ai search has become a board-level question, and why most answers to it collapse under scrutiny.
The metric itself is simple to define: the proportion of brand mentions a brand earns across a defined set of prompts and engines, relative to all brands named in those same answers. Calculating it defensibly is not. Share of voice in AI search is a procedure β a frozen prompt set, repeated runs, position-weighted mentions, and a trend line β not a single equation. Most explainers publish one formula and stop.
Alef, an AI visibility engine, tracks mentions, citations, rankings, and competitor share across ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews β the exact inputs a share-of-voice calculation consumes, alongside the AI visibility metrics that predict pipeline.
One caveat shapes everything that follows: identical prompts produce different answers across runs. A single snapshot is a sample of one, not a measurement. By the end of this guide, the reader holds a figure they can defend line by line in a meeting.
When You Need an AI Share-of-Voice Number
Four situations reliably force the question. Each one has a different urgency profile, but all four converge on the same requirement: a scored, repeatable baseline rather than a screenshot.
- A stakeholder asks for the AI visibility line item. A board deck, a QBR, or a client review includes a slide for AI search, and no one can produce a number that survives scrutiny. Alef's benchmarking research documents this as the most common trigger β the request arrives before any measurement framework exists.
- Organic click-through rates slide as AI summaries absorb queries. Pew Research Center found that Google users clicked a result in 8% of visits when an AI summary appeared, versus 15% without one β a gap that turns AI visibility from a curiosity into a revenue exposure.
- Competitors appear across every prompt run while the brand does not. This stays anecdotal until it is standardized into a scored baseline; the mechanics behind that pattern are covered in why a brand stays invisible in AI answers.
- AI-referred traffic already appears in analytics. Once sessions from ChatGPT or Perplexity show up in a report, the channel is live and measurable β AI-referred traffic measurement explains how to isolate it.
Prerequisites Before You Start
Three inputs are non-negotiable: an indexable site, a defined category, and at least three named competitors. A pre-launch site with no crawlable content has nothing for AI engines to cite, and share of voice is premature.
Budget two to four hours for a first defensible baseline. No engineering work is required, but a spreadsheet or a tracking platform is mandatory β manual spot checks will not hold up.
How to Calculate AI Share of Voice: 12 Steps
The calculation itself is not difficult. What separates a defensible AI share-of-voice figure from a misleading one is the sequence of decisions made before any number is produced: which prompts are asked, how many times each is run, which rivals count toward the denominator, and how a mention buried in the fifth paragraph is weighted against one in the opening sentence. The twelve steps below follow that order deliberately, because a flaw introduced at step two cannot be repaired at step eleven.
Time required: four to six hours for the initial build, then roughly 90 minutes per reporting cycle once the prompt set is frozen. Skill level: intermediate. Comfort with spreadsheet formulas and a basic understanding of how retrieval-augmented generation assembles an answer are sufficient. Prerequisites: a defined product category, a named competitor cohort, and access to the engines being measured. No prior share-of-voice reporting is required.
1. Define the category question, not the brand
Every prompt in the set must be phrased the way a buyer would phrase it before they know which vendors exist. A prompt such as "Is Alef a good AI visibility platform?" is worthless for measurement purposes, because naming the brand primes the model to discuss it. The answer will almost certainly contain the brand, the mention rate will look healthy, and the figure will describe nothing about actual discovery.
The correct construction removes the brand entirely: "What tools help marketing teams track how often their brand appears in AI-generated answers?" That prompt can be answered by any vendor in the category, which is precisely what makes it measurable. The brand either earns its place in the response or it does not.
This distinction matters more in AI search than it did in ranked search results. Pew Research Center found that Google users presented with an AI summary are substantially less likely to click through to any underlying link, which means the answer text itself is now the surface where a brand is either present or absent. There is no position-three listing to fall back on.
A useful test before a prompt enters the set: if the brand name were swapped for a competitor's, would the prompt still make sense as a buyer question? If not, it belongs in the brand-specific bucket at step two, not in the discovery set.
2. Build a fixed prompt set of 20 to 50 prompts across four intent buckets
The prompt set is the sample frame. Change it between reporting periods and the trend line becomes meaningless, because the underlying population has shifted. Freeze the set before the first run and treat additions as a versioned change, not an edit.
Twenty to fifty prompts is the practical range. Below twenty, a single unusual answer swings the percentage by several points. Above fifty, manual review becomes expensive enough that teams quietly stop doing it, which is worse than a smaller set run consistently.
Distribute the prompts across four intent buckets, because brands rarely perform evenly across them:
| Intent bucket | Example prompt | Typical share of set |
|---|---|---|
| Category discovery | "What are the leading platforms for monitoring brand presence in AI answers?" | 40% |
| Comparison | "How do AI visibility platforms compare on competitor tracking?" | 25% |
| Problem and solution | "How can a marketing team find out why ChatGPT never mentions their brand?" | 20% |
| Brand-specific | "What does Alef do, and who is it for?" | 15% |
The brand-specific bucket is the control. A brand that scores well there and poorly in category discovery has a positioning problem, not a visibility problem, and the two require different responses. Reporting a single blended figure hides that distinction entirely.
3. Lock the competitor cohort to three to five named rivals
Share of voice is a ratio, and a ratio needs a fixed denominator. If the competitor list grows from four to seven between Q1 and Q2, the brand's share will fall even if its absolute mention count rose, because three new brands are now absorbing mentions.
Three to five named rivals is the workable range. Fewer than three produces a figure that flatters almost any brand; more than five makes the extraction step at step seven slow and error-prone without adding much analytical value.
The cohort should be chosen once, on the basis of who actually appears alongside the brand in AI answers, and then held. A useful diagnostic: run ten category-discovery prompts manually and record which brands appear. The names that recur are the real cohort. Names that appear once are noise.
Two practical rules keep the cohort stable across reporting periods:
- Named brands only. "Other tools" or "various platforms" cannot be counted, because there is no way to attribute a mention to a specific competitor.
- Document substitutions. If a rival is acquired or exits the category, record the change and the date, and annotate the trend line rather than silently swapping the name.
4. Choose the engines deliberately and score them separately
ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews do not behave alike. They draw on different retrieval pipelines, cite different source types, and surface different brands for the same prompt. Blending them into one headline number produces a figure that describes no engine accurately.
Search Engine Journal's analysis of crawl data found that ChatGPT's crawler and Googlebot retrieve and prioritize content differently, which is one mechanical reason the same page can be cited by one system and ignored by another. The practical consequence for measurement is that each engine needs its own row in the report.
A defensible reporting structure looks like this:
| Engine | Prompts run | Runs per prompt | Total answers | Brand mentions | Mention rate |
|---|---|---|---|---|---|
| ChatGPT | 40 | 5 | 200 | 58 | 29.0% |
| Perplexity | 40 | 5 | 200 | 71 | 35.5% |
| Gemini | 40 | 5 | 200 | 44 | 22.0% |
| Copilot | 40 | 5 | 200 | 31 | 15.5% |
| Google AI Overviews | 40 | 5 | 200 | 49 | 24.5% |
A single blended figure across those five rows would read 25.3% and would obscure a fourteen-point spread between the strongest and weakest engine. That spread is the actionable finding. It tells a team where to concentrate content and citation work, and it disappears the moment the engines are averaged together.
5. Run each prompt at least three to five times
This is the step most published formulas omit, and it is the one that most often invalidates a figure in review. Large language models are non-deterministic. The same prompt submitted twice can return different brands, different orderings, and different citation sets.
The assumption that lowering the sampling temperature to zero eliminates this variance is incorrect. Research published on arXiv has shown that temperature-zero decoding does not guarantee identical outputs across runs, because of factors including floating-point non-associativity and batch-level inference effects. A single run per prompt therefore samples one path through a distribution and reports it as if it were the distribution.
Three to five runs per prompt is the practical minimum. At forty prompts and five runs, the sample is 200 answers per engine, which is large enough that a single anomalous response moves the mention rate by half a percentage point rather than several. That stability is what makes the number defensible when someone asks how it was produced.
Two implementation notes:
- Vary the run conditions. Where the engine supports it, run some prompts in a fresh session and some in an existing one, since conversation context can influence retrieval.
- Record every run, including the boring ones. Discarding runs that "look wrong" is a form of selection bias and will inflate the figure.
6. Capture the raw answer text and timestamp every run
An AI share-of-voice figure that cannot be audited is an assertion, not a measurement. Every run should be stored with the full answer text, the engine, the prompt as submitted, the run number, and a timestamp.
This serves three purposes. It allows a disputed mention to be checked directly rather than re-litigated from memory. It allows the extraction logic at step seven to be re-run against the same corpus if the rules change. And it makes the trend line at step twelve comparable, because the underlying evidence for each period still exists.
Storage format matters less than completeness. A structured table with one row per answer is sufficient:
| Field | Example value |
|---|---|
| Run ID | CHATGPT-0042-R3 |
| Engine | ChatGPT |
| Prompt ID | CD-017 |
| Run number | 3 |
| Timestamp (UTC) | 2025-11-14T09:22:41Z |
| Answer text | Full response, unedited |
| Citations returned | List of URLs, if any |
The citation field is worth capturing even when the metric does not use it. Semrush's ghost citations research examined the gap between sources that AI systems cite and sources that receive measurable referral traffic, and that gap is only visible if citations are recorded at the point of collection rather than reconstructed later.
7. Extract every brand mention and record the order of appearance
Extraction is where a clean sample becomes a dataset. For each stored answer, identify every brand from the locked cohort that appears, plus the brand being measured, and record the position at which each appears.
Position is not cosmetic. An answer that opens with a brand and then lists four alternatives is not equivalent to one that mentions the brand in a closing caveat. Recording ordinal position β first, second, third, or later β preserves that distinction for the weighting step.
Three extraction rules prevent the most common errors:
- Count the brand once per answer, not once per sentence. A brand named three times in one response is one mention for mention-rate purposes, with the first occurrence determining position.
- Match on the brand and its known variants. Product names, former names, and common abbreviations should be mapped to a single canonical entity before counting.
- Log negative and neutral mentions separately. A brand described as unsuitable for a use case is present in the answer but is not earning the same value as a recommendation. Tracking sentiment alongside presence prevents a mention rate from overstating competitive position.
For teams running this across multiple engines and hundreds of answers, the extraction step is the first place manual workflows break down. Alef's prompt intelligence workspace handles prompt-set management and mention extraction across engines in one place, which removes the copy-paste layer that typically introduces errors here. A more detailed walkthrough of the extraction mechanics is available in this guide to tracking brand mentions in ChatGPT and Perplexity.
8. Choose between mention rate and share of mentions
Two metrics are routinely conflated under the label "share of voice," and they answer different questions.
Mention rate is the brand's mentions divided by the total number of answers. If the brand appears in 58 of 200 ChatGPT answers, the mention rate is 29.0%. This metric can exceed 100% when summed across brands, because a single answer can mention five competitors. It measures presence: how often the brand shows up at all.
Share of mentions is the brand's mentions divided by all brand mentions recorded across the cohort. If the brand appears 58 times and the cohort's total mentions across all answers is 310, the share of mentions is 18.7%. This metric sums to exactly 100% across the cohort. It measures competitive position: how much of the category conversation the brand occupies.
| Metric | Formula | Sums to 100%? | Answers the question |
|---|---|---|---|
| Mention rate | Brand mentions Γ· total answers | No | How often is the brand present? |
| Share of mentions | Brand mentions Γ· all cohort mentions | Yes | How much of the category does the brand hold? |
The choice depends on the audience. A content team tracking whether new pages are being picked up will find mention rate more responsive, because it moves when coverage expands. A leadership team comparing the brand against three named rivals will find share of mentions more legible, because it behaves like market share and sums to a whole.
Reporting both, side by side, is the safest option. The two figures together distinguish a brand that is growing in absolute presence from one that is growing only because the category is expanding.
9. Weight mentions by position in the answer
A flat mention count treats every appearance as equal. Buyers do not. The brand named first in a recommendation list receives attention that a brand named fifth does not, and the measurement should reflect that asymmetry.
A simple, defensible weighting scheme assigns a multiplier based on ordinal position:
| Position in answer | Weight |
|---|---|
| First brand mentioned | 1.5 |
| Second | 1.25 |
| Third | 1.0 |
| Fourth or later | 0.75 |
Applying these weights converts the raw mention count into a weighted score. A brand with 58 mentions, of which 20 were first-position and 38 were third-or-later, produces a weighted score of 68.0 rather than 58.0. Run the same calculation for each competitor and the denominator becomes the sum of weighted scores, which yields a position-weighted share of voice.
The weights themselves are a judgment call, and any reasonable scheme is defensible provided it is applied consistently and disclosed. What is not defensible is changing the weights mid-year without restating prior periods, because that breaks comparability in exactly the way a shifting prompt set does.
10. Decide how to treat citations separately from mentions
A brand can be mentioned in an answer without being cited as a source, and it can be cited without being named in the prose. These are different outcomes and should be counted separately.
The IAB's framework for measuring visibility in the AI era treats presence and attribution as distinct dimensions, which is the right structure for reporting. A mention indicates the model considers the brand relevant to the question. A citation indicates the model drew on the brand's own content to construct the answer. The second is closer to the traditional organic-search outcome and tends to correlate with downstream referral traffic.
Report them as two columns rather than merging them into one score. A brand with a 29% mention rate and a 6% citation rate has a content problem: it is known but not used as a source. A brand with a 12% mention rate and an 11% citation rate has the opposite profile and a different set of next actions.
11. Normalize across engines before comparing periods
Because each engine is scored separately at step four, any cross-engine or period-over-period comparison needs a consistent basis. The cleanest approach is to report each engine's figure against the same prompt set and the same run count, then present a simple unweighted average of the engine-level figures as a secondary line, clearly labeled as an average rather than a blended total.
If an engine is added mid-year β a new assistant enters the market, for instance β do not fold it into the historical series. Report it as a new row with its own start date. Retroactively blending a new engine into prior periods implies data that was never collected.
12. Trend the figure over time rather than reading a single snapshot
A single AI share-of-voice number is close to meaningless on its own. It has no baseline, no variance, and no direction. The figure becomes useful only when the same frozen prompt set is run on a regular cadence β monthly is typical, quarterly is acceptable β and the results are plotted.
Three things become visible in a trend line that a snapshot cannot show:
- Whether a content initiative moved the needle. A rise from 18.7% to 24.1% over two quarters, against an unchanged prompt set and cohort, is attributable in a way that a single reading never is.
- Whether a competitor's gain came at the brand's expense. Share of mentions sums to 100%, so movement in one brand's line is necessarily reflected elsewhere in the cohort.
- Whether variance is noise or signal. With five runs per prompt, the run-to-run standard deviation establishes a band. Movement inside that band should not be reported as a trend.
Set the cadence, hold the prompt set and cohort fixed, and annotate every change to either. The resulting series is the artifact that survives scrutiny in a meeting, because every point on it can be traced back to stored answer text with a timestamp.

Common Mistakes That Invalidate an AI Share-of-Voice Figure
Most inflated or indefensible AI share-of-voice numbers fail for the same handful of reasons. Each is avoidable, and each fix takes minutes rather than days.
The Eight Failures That Break the Metric
- Single-run snapshots. Reading one ChatGPT answer and calling it a measurement ignores non-determinism entirely; run each prompt at least five times and average the results.
- Brand-primed prompts. Prompts that name the brand inflate the share and destroy comparability; keep the prompt set unbranded so every competitor competes on equal footing.
- Blended engines. Averaging ChatGPT and Perplexity into one figure hides where visibility is strong or weak; report per-engine shares, then aggregate.
- Confusing mention rate with share of mentions. Mention rate can exceed 100% across brands, while share of mentions sums to 100%; mixing the two produces a meaningless figure.
- Ignoring position. Treating a fourth-place mention as equal to a first-place recommendation overstates prominence; apply position weights before aggregating.
- Counting citations as mentions. Semrush found that 61.7% of citations are ghost citations that never name the brand, which is why citation tracking has to separate mentions from links.
- Changing the prompt set mid-quarter without logging it. Silent edits make trend lines uninterpretable; version the prompt set and record every change.
- Treating AI share of voice as a Google ranking proxy. The two move independently, as AI visibility and Google rankings track different signals; blending them obscures both.
Any one of these errors is enough to turn a defensible figure into a talking point that collapses under questioning.
Summary Table: Steps, Inputs, and Expected Outcomes
The twelve steps below form one repeatable procedure. Each row names the action, the metric it produces, and the check that confirms the step was executed correctly β so a reviewer can audit the calculation without re-running it. Teams benchmarking these figures against rival brands can extend the same structure using the competitor AI visibility benchmarking framework.
| Step | What You Do | Metric Produced | Verification Check |
|---|---|---|---|
| 1 | Define the prompt set from real buyer questions | Frozen prompt list (50β200 prompts) | Every prompt maps to a purchase-stage query |
| 2 | Freeze prompt wording, engine, and locale | Versioned prompt file (v1.0, dated) | Two analysts export identical prompt text |
| 3 | Run each prompt 3β5 times per engine | Raw response log | Variance across runs under 10% |
| 4 | Extract every brand mention per response | Mention count per prompt | Two coders agree on 95% of mentions |
| 5 | Classify mentions as owned, earned, or competitor | Mention type tally | No unclassified mentions remain |
| 6 | Record position of each mention in the answer | Average mention position (1st, 2nd, 3rd) | Position logged for 100% of mentions |
| 7 | Choose mention rate or share of mentions | Primary AI share-of-voice figure | Formula documented before reporting |
| 8 | Apply position weights to each mention | Weighted share of voice | Weighted total reconciles to raw count |
| 9 | Capture citation links separately from mentions | AI citation share | Cited URLs resolve to live pages |
| 10 | Aggregate across engines into one view | Cross-engine share of voice | Engine-level sums match the total |
| 11 | Compare against 3β5 named competitors | Relative share delta | Competitor set unchanged since v1.0 |
| 12 | Re-run monthly and plot the trend | 90-day share-of-voice trend line | At least three data points plotted |
A single snapshot answers "where do we stand today." Only the trend line answers "is the number moving" β and that is the version worth defending in a meeting.

Conclusion: The Figure You Can Defend
A defensible share-of-voice number is a procedure, not a formula: a frozen prompt set, repeated runs, position-weighted mentions, and a trend line that shows direction rather than a single snapshot. Mention rate answers whether the brand appears at all; share of mentions answers how it ranks against competitors. Choose based on the question being asked, not on which number looks better.
Non-determinism is why repeated runs are non-negotiable β the same prompt returns different answers across sessions, so one execution proves nothing. As answer surfaces continue shifting, tracking those movements over time is what keeps the metric honest, a dynamic Alef's roundup of AI answer engine trends reshaping brand visibility documents in detail.
Key takeaways - A defensible figure comes from a repeatable procedure: frozen prompts, multiple runs, position weighting, trend tracking. - Mention rate measures presence; share of mentions measures competitive standing. - Non-deterministic outputs make single-run measurements unreliable by default. - One snapshot is a data point; a trend line is a defensible number.
Frequently Asked Questions
What is share of voice in AI search?
Share of voice in AI search is the percentage of answer slots across a defined prompt set in which a brand is mentioned, relative to all brands mentioned in those same answers. The numerator counts the brand's appearances β every mention, citation, or recommendation the model surfaces for the tracked prompts. The denominator counts all brand appearances across that identical prompt set, which makes the metric zero-sum: when one brand gains a mention, another loses one. Because the denominator is bounded by the prompt set itself, an AI share-of-voice figure is only as meaningful as the prompt list it was measured against.
How do you calculate AI share of voice?
The calculation divides a brand's total weighted mentions by all brands' weighted mentions across the same frozen prompt set, then multiplies by 100. The critical requirement is repeated runs: because large language models are non-deterministic, a single pass produces a snapshot, not a measurement. Each prompt must be executed multiple times, with every run's mentions extracted and aggregated before the ratio is computed. Teams that skip repetition end up defending a number that shifts the moment anyone re-runs the query.
What is the difference between mention rate and share of mentions?
Mention rate measures how often a brand appears at all β the percentage of runs in which it is named β while share of mentions measures its portion of all brand mentions in those runs. The distinction is zero-sum versus non-zero-sum. Mention rate can rise for every competitor simultaneously if models simply list more brands per answer, but share of mentions cannot: gains come only at rivals' expense. For competitive reporting, share of mentions is the harder, more defensible figure, which is why AI visibility reporting frameworks treat the two as separate dashboard metrics rather than substitutes.
Why do AI answers change between runs?
AI answers change between runs because model outputs are probabilistic, and even a temperature setting of zero does not guarantee identical responses across executions. Research on LLM determinism has documented that identical prompts can yield different outputs due to factors such as floating-point non-associativity and hardware or batch-size variation. This is why any share-of-voice methodology built on a single query execution is structurally unreliable β the variance is a property of the system, not a flaw in the prompt.
How many times should you run each prompt?
At least three to five runs per prompt is the practical minimum, with variance monitored across runs rather than averaged away silently. If a brand appears in two of five runs, the honest figure is 40 percent mention rate with a wide confidence band β not "sometimes mentioned." Tracking the spread matters as much as the mean, since a metric that swings wildly between runs cannot anchor a quarterly trend line. A platform that logs every run preserves that variance; manual spot checks typically discard it.
Can you track brand share of voice in ChatGPT specifically?
Yes β share of voice can be scored per engine, and per-engine scoring is more useful than a blended figure. ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews draw on different retrieval and ranking behaviors, so a brand can hold strong share in one engine and near-zero in another. Alef's workspace computes mention, citation, and competitor share across each of those engines separately, which makes it possible to see where the gap actually sits β a level of granularity that AI visibility tracking requires to be actionable. Blending engines into one number hides exactly the signal a marketing lead needs in the meeting.
Start Measuring Your AI Share of Voice
A defensible baseline starts with a single workspace rather than a spreadsheet. Alef's AI visibility platform tracks brand mentions, citations, rankings, and competitor share across ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews, applying the same repeated-run and position-weighted logic described above. The first baseline typically takes an afternoon to configure. Start at alef.ink and bring the resulting figure to the next meeting β with the method behind it documented.
Sources
- Pew Research Center β Google users are less likely to click on links when an AI summary appears
- Semrush β The Ghost Citations Study
- Search Engine Journal β ChatGPT vs Googlebot crawl data
- IAB β Measuring Visibility in the AI Era
- arXiv β An Empirical Study of the Non-determinism of ChatGPT in Code Generation
- arXiv β Who Owns the AI Recommendation? A Multi-Industry Empirical Map of Brand Category Ownership Across Large Language Models
- Search Engine Land β The ghost citation problem in AI answers
Turn this article into a visibility plan
Use Alef to audit your site, find content gaps, and create briefs your team can ship.