How to Benchmark AI Visibility Against Competitors in 5 Steps (With a Repeatable Scoring Model)
Learn how to benchmark AI visibility against 3-5 competitors: build a fixed prompt set, sample repeated runs, and score mentions and citations separately.

Intro
ChatGPT's crawler now issues roughly 3.6 times more requests to websites than Googlebot, according to a Search Engine Journal analysis of 24.4 million proxy requests across 69 sites β yet many marketing teams still benchmark AI visibility with a single Google rank check. Consider the scenario: a buyer asks ChatGPT for the best vendor in the category, and the answer names three competitors but not the brand. Where in any existing dashboard does that loss appear?
A repeatable AI visibility benchmark answers that question. It is a fixed prompt set run repeatedly against a locked competitor set, scored on mentions and citations separately, and reviewed on a schedule. The outcome is a defensible baseline showing where the brand stands versus three to five named competitors across every tracked engine.
Alef's AI visibility solution tracks mentions, citations, share of voice, and competitor context across ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews, so the procedure described here mirrors what the platform runs continuously. Existing share-of-voice explainers define the formula; this guide supplies the operating procedure they omit.
When You Need a Competitive AI Visibility Benchmark
The trigger is rarely a strategic planning session. It is a screenshot. A buyer runs a high-intent prompt in ChatGPT or Perplexity, the answer names three competitors and omits the brand entirely, and the request that lands on the marketing lead is blunt: prove where we stand.
Four signals typically precede that moment:
- Organic click-through rates decline as AI summaries absorb queries. Pew Research Center found that when an AI Overview appeared in March 2025, users clicked a traditional result in only 8% of visits, versus 15% without a summary (Pew Research Center).
- Competitors appear consistently across repeated prompt runs while the brand does not β anecdotal until standardized into a scored baseline. When a brand is invisible in AI answers, the cause is usually structural, not random.
- AI-referred traffic already appears in analytics, confirming the channel is live and measurable. Distinguishing it from other referral sources requires knowing how AI-referred traffic is defined and measured.
- A stakeholder, board, or client requests an AI visibility line item in a quarterly review, and no defensible number exists.
A benchmark is premature for pre-launch sites with no indexable content. Fix indexation and crawlability first; answer engines cannot cite what they cannot retrieve.
The 5-Step Benchmark Process (10 Concrete Actions)
A benchmark is only as defensible as the procedure that produced it. The ten actions below build a competitive AI visibility baseline that can survive scrutiny from a CFO, a board, or a skeptical agency partner β because every number traces back to a frozen prompt set, a locked competitor list, and a logged sampling method.
Prerequisites and Time Commitment
Before running the first benchmark, four conditions need to be in place:
- Intermediate SEO or go-to-market knowledge. The person running the benchmark should understand query intent classification, branded versus non-branded demand, and how answer engines assemble sources. This is not a task to hand to someone who has never built a keyword map.
- A published, crawlable site. If the brand's own domain is blocked by robots.txt rules aimed at AI crawlers, or if key pages are orphaned, citation scoring will understate reality. The crawl landscape has shifted dramatically β Search Engine Journal's analysis of 24 million requests found ChatGPT now crawls 3.6x more than Googlebot, which means crawl accessibility is now a direct input to AI visibility.
- A defined category. "We sell software" is not a category. "We sell revenue intelligence software to mid-market B2B SaaS companies" is. Prompts cannot be written against an undefined category.
- A realistic time budget. The first benchmark takes roughly 6 to 10 hours: prompt construction, competitor selection, engine runs, and scoring. Monthly refreshes run 2 to 3 hours once the scaffolding exists.
The output of these ten actions is a scored baseline: mention rate, citation rate, share of voice, and position/sentiment notes for the brand and every locked competitor, per engine.
Step 1 β Define the Prompt Set
Build 40 to 60 buyer-intent prompts across four buckets, written in the buyer's phrasing.
The prompt set is the measuring instrument. If it is sloppy, every downstream number inherits the error. Four buckets keep coverage balanced:
- Category discovery β "What are the best AI visibility platforms for B2B SaaS?" These capture buyers who know the problem space but not the vendors.
- Comparison β "How does [Brand A] compare to [Brand B] for enterprise SEO tracking?" These capture late-stage evaluation and are where share of voice moves fastest.
- Problem-solution β "How do I track whether ChatGPT mentions my brand?" These capture buyers describing a pain point without vendor vocabulary.
- Brand-adjacent β "Is [Your Brand] worth it?" or "What do people say about [Your Brand]?" These measure reputation framing and are frequently the most uncomfortable bucket to read.
Write prompts the way a buyer types them into ChatGPT or Perplexity, not the way they appear in a keyword tool. "AI visibility measurement" is a keyword. "How do I know if AI search engines are recommending my company?" is a prompt. The distinction matters because answer engines respond to natural language framing, and keyword-shaped prompts produce keyword-shaped answers that do not reflect real buyer behavior.
Aim for 10 to 15 prompts per bucket. Once the set is built, freeze it. Changing prompts mid-stream invalidates month-over-month deltas β a share-of-voice increase that comes from swapping three prompts is not a share-of-voice increase. Alef's prompt intelligence workflow is built around maintaining a stable, versioned prompt library for exactly this reason: the value of the benchmark comes from comparability across time, not from novelty in the prompt set.
Expected outcome: A versioned prompt file (spreadsheet or database) with 40 to 60 prompts, each tagged by bucket, intent stage, and target persona. Version-stamp it with a date. That date is the benchmark's epoch.
Step 2 β Lock the Competitor Set
Choose three to five competitors that appear in the same buyer consideration set, and freeze the list for at least two quarters.
Three to five is the practical range. Fewer than three produces a share-of-voice denominator too small to be stable; more than five dilutes the analysis and multiplies scoring time without adding decision value.
The selection criterion is not "who has the biggest market share." It is "who shows up when our buyers ask an answer engine for options." Those are often different lists. A useful test: run five category-discovery prompts and record every brand the engines name. The three to five that recur most often are the real consideration set, regardless of how the org chart feels about them.
Freeze the list for a minimum of two quarters. Share-of-voice math depends on a constant denominator β if a competitor is swapped in mid-year, every historical comparison becomes apples-to-oranges. If a genuinely new entrant starts appearing in answers, log it as an observation and add it at the next scheduled reset, not mid-cycle.
Expected outcome: A locked competitor roster with a documented rationale for each inclusion, plus a scheduled review date two quarters out.
Step 3 β Select the Engines and Record the Configuration
Track ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews where relevant, and log model version, date, and locale for every run.
Engine selection should follow the audience. A B2B SaaS brand selling to US marketing teams will care most about ChatGPT, Perplexity, and Google AI Overviews. A consumer brand may weight Gemini and Copilot more heavily. There is no universal answer, but there is a universal requirement: whatever engines are selected, the configuration must be recorded.
Record for every run:
- Engine and model version (for example, GPT-4o versus a successor release)
- Date and time of the run
- Locale and language settings
- Whether the run was logged in or logged out
- Whether personalization or memory features were active
Answers shift as models update. A benchmark run in March against one model version is not directly comparable to a June run against a newer one β and without the configuration log, that shift looks like a visibility change when it is actually a model change. The configuration log is what allows a team to say, credibly, "our share of voice moved for a reason we can name."
Expected outcome: A configuration log with one row per engine per run date, stored alongside the prompt file.
Step 4 β Run Repeated Samples per Prompt to Absorb Non-Determinism
Run each prompt three to five times per engine and score the distribution rather than a single answer.
This is the step most homegrown benchmarks skip, and it is the step that most often invalidates them. Large language models are non-deterministic: the same prompt against the same model can produce materially different answers across runs. Research on hosted environments has documented that even "deterministic" system settings β temperature zero, fixed seed β do not guarantee identical outputs across repeated calls, a phenomenon examined in detail in ACL Anthology's study of non-determinism in hosted LLM settings. Related work on non-deterministic drift in large language models has reported accuracy swings of up to 15% across repeated runs of identical inputs.
A 15% swing is larger than most quarter-over-quarter visibility changes. A benchmark built on one sample per prompt is measuring noise and calling it signal.
The fix is straightforward: run each prompt three to five times per engine, then score the distribution. If a brand appears in 4 of 5 runs, mention rate for that prompt is 80%. If it appears in 1 of 5, it is 20%. The aggregate across the prompt set is the benchmark number. Three runs is the floor; five is better for high-stakes prompts in the comparison bucket, where a single answer can swing a deal.
Expected outcome: A raw results matrix with one row per prompt-run and columns for engine, run number, brands mentioned, brands cited, and position notes. The matrix is the source of truth; every score is derived from it.
Step 5 β Score Mentions and Citations Separately
Mention rate measures whether the brand name appears in the answer text; citation rate measures whether the brand's own domain is listed as a source.
These two metrics are frequently conflated, and conflating them hides the most actionable finding in the entire benchmark. A brand can be mentioned without being cited β the engine names it from training data but links elsewhere. A brand can be cited without being mentioned β the domain appears in a source list but the answer text never names it. Each failure mode has a different remedy.
The gap between mention and citation varies sharply by engine. A 543-query study cited in industry research found that Perplexity cited a source in 60.6% of answers, versus 32.6% for ChatGPT. That divergence has a direct operational implication: a citation-building strategy that works on ChatGPT may be under- or over-indexed relative to Perplexity, and the benchmark is what reveals which.
Score both metrics per prompt, per engine, per run, then aggregate:
| Metric | Definition | What a low score means |
|---|---|---|
| Mention rate | Share of sampled answers where the brand name appears in the answer text | The model does not associate the brand with the category |
| Citation rate | Share of sampled answers where the brand's own domain appears as a cited source | The brand's content is not being retrieved or is not citation-worthy |
For the underlying metric definitions and how they map to revenue outcomes, the cluster pillar How to Measure AI Visibility: The Metrics That Predict Revenue, Not Just Mentions covers the scoring model in full. This section assumes those definitions and applies them competitively.
Expected outcome: Two separate scores per brand per engine β mention rate and citation rate β with the gap between them documented as a finding in its own right.
Step 6 β Compute Share of Voice Against the Locked Set
Divide brand mentions by total answers measured, reported per engine.
Share of voice answers the question a single mention rate cannot: how does the brand's presence compare to the competitive field? The formula is brand mentions divided by total answers measured, computed per engine and per bucket.
One structural detail trips up most first-time benchmarkers: multiple brands can appear in a single answer. A category-discovery prompt may name five vendors. This means share-of-voice figures across the locked set can sum above 100%. That is not an error β it reflects the reality that answer engines present option sets rather than single recommendations. The metric is still useful for tracking relative position over time, provided the same caveat is stated wherever the number appears.
Report share of voice per engine, never as a single blended figure. A brand at 40% share of voice on ChatGPT and 12% on Perplexity has a very different strategic situation than one at 26% on both, even though the averages are identical.
Expected outcome: A share-of-voice table with brands as rows, engines as columns, and a footnote stating that shares may exceed 100% due to multi-brand answers.
Step 7 β Capture Position and Sentiment, Not Just Presence
Note whether the brand is named first, in a shortlist, or as an also-ran β and whether the framing is positive, neutral, or cautionary.
Presence is binary; position is ordinal; sentiment is qualitative. All three matter, and the third is the one that generates the most uncomfortable but most valuable findings.
Position categories worth tracking:
- First-named β the brand leads the answer
- Shortlisted β the brand appears in a set of options without leading
- Also-ran β the brand appears late, in a caveat, or in a "other options include" tail
- Absent β the brand does not appear
Sentiment categories:
- Positive β the answer frames the brand favorably, with specific strengths
- Neutral β the answer names the brand without evaluation
- Cautionary β the answer includes limitations, caveats, or comparisons that position the brand as a weaker option
A brand with a 70% mention rate where two-thirds of mentions are cautionary is in worse shape than a brand with a 40% mention rate where nearly every mention is positive and first-named. Mention rate alone would rank these backwards. Position and sentiment scoring corrects the ranking and points directly at the content or reputation work that needs to happen.
Scoring sentiment requires judgment, so document the rubric before scoring begins and apply it consistently. A three-point scale with written examples per point is sufficient; a ten-point scale invites drift between scorers.
Expected outcome: A position and sentiment column added to the results matrix, with the scoring rubric saved alongside the prompt file so future refreshes apply the same standard.
Step 8 β Normalize the Scores into a Single Benchmark Index
Convert mention rate, citation rate, share of voice, and position/sentiment into a weighted index so the brand and each competitor can be ranked on one number.
Four separate metrics are informative but hard to act on. A single index β with the component weights documented β makes the benchmark decision-ready. A defensible default weighting for a consideration-stage benchmark:
| Component | Weight | Rationale |
|---|---|---|
| Mention rate | 30% | Baseline presence in answers |
| Citation rate | 30% | Direct evidence the brand's own domain is being retrieved |
| Share of voice | 25% | Relative position against the locked competitor set |
| Position and sentiment | 15% | Quality of presence, not just quantity |
The weights are a judgment call and should reflect the business model. A brand whose revenue depends on AI-referred traffic should weight citation rate higher, because citation is the mechanism by which a domain earns the click. A brand in a crowded category where buyers shortlist from answers should weight share of voice higher. The point is not the specific weights β it is that the weights are declared before scoring, so the index cannot be reverse-engineered to produce a flattering result.
Publish the index with its component scores visible. An index that hides its inputs is not a benchmark; it is an assertion.
Expected outcome: A single benchmark index per brand per engine, with component scores and weights shown in the same table.
Step 9 β Set the Review Cadence and Change Log
Refresh monthly, re-baseline quarterly, and log every change to prompts, competitors, or configuration.
Cadence is what separates a benchmark from a one-off audit. A monthly refresh at 2 to 3 hours keeps the data current enough to catch meaningful movement without consuming the team's capacity. A quarterly re-baseline allows for deliberate changes β adding a new competitor, retiring prompts that no longer reflect buyer language, updating engine model versions β while preserving comparability within each quarter.
Every change goes in a change log with a date and a reason. When share of voice moves 8 points between February and March, the change log is what distinguishes a real market shift from a prompt that was quietly edited. Without it, the benchmark loses its defensibility the first time someone asks a hard question about methodology.
Expected outcome: A cadence calendar with monthly refresh dates and quarterly re-baseline dates, plus a change log with one row per modification.
Step 10 β Distribute the Benchmark as a Standing Report
Package the results into a one-page report with the index, component scores, and the three biggest gaps β and route it to the people who can act on it.
A benchmark that lives in a spreadsheet only its author opens has no organizational effect. The report format that works is short: one page, index scores for the brand and each competitor, the component breakdown, and a short list of the largest gaps β for example, "citation rate on Perplexity is 14 points below the competitive median" or "brand is absent from 60% of comparison-bucket prompts."
Route it to content, SEO, and product marketing on a fixed schedule. The gaps in the benchmark map directly to work: citation gaps point at content retrieval and citation-worthiness, mention gaps point at category association, and sentiment gaps point at reputation and comparison content. Alef's visibility engine runs this measurement loop continuously rather than monthly, which is the natural next step once a manual benchmark has established the baseline and proven the value of the exercise.
Expected outcome: A recurring one-page benchmark report with a named owner, a distribution list, and a standing agenda item where the gaps are converted into assigned work.
Common Mistakes That Invalidate an AI Visibility Benchmark
Most failed benchmarks collapse for the same handful of reasons, and each is preventable with a small amount of upfront discipline.
Single-Sample Scoring
One run per prompt produces a snapshot, not a benchmark. Hosted large language models return different answers across identical requests, a drift documented in research on non-determinism in hosted LLM environments. Average three to five runs per prompt, per engine, before drawing any conclusion.
Changing the Prompt Set Mid-Cycle
Adding or rewording prompts between measurement periods makes month-over-month deltas meaningless. Freeze the set, version it explicitly, and log every change with a date.
Conflating Mentions With Citations
A brand can be named in an answer while its domain is never cited, and citation rates vary by engine. Reporting one blended number hides the actual gap. Tracking brand mentions across ChatGPT and Perplexity as separate signals keeps the diagnosis honest.
Comparing Against an Unstable Competitor List
Swapping competitors between periods resets share-of-voice math to zero and destroys trend comparability. Lock the list for at least two full cycles.
Ignoring Engine and Model Version
Answers shift as models update, so a benchmark without a recorded model version and date cannot be reproduced or defended.
Treating the Benchmark as a Report Rather Than a Queue
A benchmark that does not end in a prioritized list of prompts to fix is a vanity artifact. Before committing budget to any measurement stack, it is worth knowing how to evaluate an AI SEO platform before you buy so the output maps to action.
Skipping the Source Layer
Tracking only whether the brand appears, and not which domains the engines cite, removes the only actionable lever for earning future citations.
Benchmark Integrity Checklist
- Minimum sample size: At least three runs per prompt per engine, averaged before scoring.
- Frozen prompt set: No additions or rewording within a measurement period.
- Versioned records: Model name, version, and run date logged for every sample.
- Separate scores: Mention rate and citation rate reported independently.
- Stable competitor list: Unchanged across at least two consecutive cycles.
- Actionable output: Every benchmark closes with a ranked list of prompts to improve.
Benchmark Summary Table: Steps, Outputs, and Cadence
The table below consolidates the full benchmark procedure into a single reference. Every row maps an action to the artifact it produces, a defensible target value, the person accountable for it, and how often it needs refreshing. For teams that need to convert these outputs into a shareable client-facing document, the process for building an AI search visibility report from benchmark data extends this structure into a formatted deliverable.
| Step | What It Produces | Concrete Target or Value | Owner | Refresh Cadence |
|---|---|---|---|---|
| 1. Define the prompt set | Frozen list of buyer-intent queries | 40 to 60 prompts | Content lead | Quarterly |
| 2. Segment prompts by intent | Tagged prompt groups (awareness, comparison, purchase) | 3 segments, minimum 10 prompts each | Content lead | Quarterly |
| 3. Select competitors | Locked competitor list | 3 to 5 named competitors | Marketing lead | Quarterly |
| 4. Choose engines | Fixed engine panel | 5 engines tracked | Marketing lead | Semi-annually |
| 5. Run repeated samples | Raw response corpus | 3 to 5 samples per prompt per engine | Analyst | Monthly |
| 6. Score mention rate | Brand appearance frequency | Reported as a percentage, separate from citations | Analyst | Monthly |
| 7. Score citation rate | Linked or attributed source frequency | Reported as a percentage, separate from mentions | Analyst | Monthly |
| 8. Calculate share of voice | Brand share across the prompt set | Sum of brand mentions divided by all tracked mentions | Analyst | Monthly |
| 9. Compare against competitors | Ranked competitive position | Rank order plus percentage-point gap to leader | Marketing lead | Monthly |
| 10. Verify reproducibility | Second-analyst confirmation | Share of voice reproduced within a few percentage points | Second analyst | On each refresh |
| Review cadence | Scheduled re-run and diff report | Monthly for fast-moving categories, quarterly otherwise | Marketing lead | Monthly or quarterly |
The verification row matters more than it appears. Because large language models drift between identical runs, a benchmark that cannot be reproduced by a second analyst is not a benchmark β it is an anecdote.
Conclusion: From Benchmark to Prioritized Action
The five-step spine holds: fix the prompt set, lock three to five competitors, sample each prompt repeatedly to absorb non-deterministic answers, score mention and citation separately, and review on a cadence. That structure turns anecdotal anxiety about AI visibility into a defensible number β one that survives stakeholder scrutiny and comparison across quarters.
Once the benchmark exposes which prompt clusters the brand loses, the work shifts from measurement to execution: closing citation gaps and producing content that earns retrieval. That is where Alef's content growth solutions connect the baseline to the fix, and the same engine tracks whether the gap closes over time.
Key takeaways - A benchmark is a fixed prompt set, a locked competitor set, and repeated samples β not a single screenshot. - Non-determinism is a sampling problem: repeat each prompt and score the distribution, not one answer. - Mentions and citations are separate scores; a brand can be named without being linked as a source. - Review monthly or quarterly so the benchmark becomes a trend line, not a one-off audit.
Frequently Asked Questions
How many prompts do I need to benchmark AI visibility?
Forty to sixty buyer-intent prompts is enough for a defensible baseline, provided they are distributed across four buckets rather than clustered around a single theme. A workable split is roughly fifteen category discovery prompts ("best [category] platforms"), fifteen comparison prompts ("[brand] vs [competitor]"), fifteen problem-solution prompts ("how to [outcome the product delivers]"), and ten brand-adjacent prompts that test whether the model associates the brand with its core use case. Fewer than forty prompts produces a sample too thin to distinguish a real visibility gap from run-to-run noise; more than sixty rarely changes the ranking of competitors and mostly adds sampling cost.
How many times should I run each prompt?
Three to five runs per prompt per engine is the practical minimum, because large language models do not return identical answers to identical inputs. Research on non-deterministic drift shows that hosted environments can produce accuracy swings of up to 15% across repeated runs of the same prompt (arXiv β Quantifying non-deterministic drift in large language models), and the ACL Anthology study on supposedly deterministic system settings reached similar conclusions (ACL Anthology β Non-Determinism of 'Deterministic' LLM System Settings in Hosted Environments). A single run per prompt measures one roll of the dice, not a position.
What is the difference between mention rate and citation rate?
Mention rate counts the answers where the brand name appears anywhere in the response; citation rate counts only the answers where the brand's own domain is listed as a linked source. The two diverge sharply by engine. In one 543-query study, Perplexity cited a source in 60.6% of answers versus 32.6% for ChatGPT, meaning the same brand can hold a strong mention position on one engine and almost no citation footprint on another. Scoring them separately is what prevents a misleading "we're visible" conclusion, as covered in how AI search visibility differs from Google rankings.
How often should an AI visibility benchmark be refreshed?
Monthly for volatile categories and quarterly for stable ones, always re-running the identical frozen prompt set so the trend line stays comparable. Changing prompts mid-cycle resets the baseline and makes month-over-month movement unreadable. For context on how quickly the underlying landscape shifts, ChatGPT now crawls roughly 3.6 times more than Googlebot across 24 million requests analyzed (Search Engine Journal), and the broader 2026 numbers are collected in these AI search statistics.
Can I benchmark AI visibility without a paid tool?
Yes, manually, using a spreadsheet and a fixed query log, but the sampling burden is what makes a tracking platform the practical choice. Forty to sixty prompts times three to five runs times five engines is 600 to 1,500 individual answer captures per cycle, each requiring a manual check for brand mention and source citation. That is a recurring monthly workload no marketing lead sustains for more than a quarter or two.
Which AI engines should be included in the benchmark?
ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews, weighted by where the target audience actually asks questions. Citation overlap between engines is low β one analysis found only about 11% of domains cited by both ChatGPT and Perplexity β so a benchmark limited to a single engine systematically overstates coverage. Tracking all five is what reveals which engine represents the largest addressable gap.
Sources
- Search Engine Journal β ChatGPT Now Crawls 3.6x More Than Googlebot: What 24M Requests Reveal
- Pew Research Center β Google users are less likely to click on links when an AI summary appears in the results
- arXiv β Quantifying non-deterministic drift in large language models
- ACL Anthology β Non-Determinism of 'Deterministic' LLM System Settings in Hosted Environments
- Zenox Media β AI Citation Study: 543 Real Answers From 4 AI Engines, Scored
- Thalox β Answer Engine Citation Overlap Strategy (ChatGPT and Perplexity share 11% of domain citations)
- Clutch β 65% of consumers use AI to research products before making a purchase
- Search Engine Land β Google's AI Overviews are hurting clicks: Pew study
Turn this article into a visibility plan
Use Alef to audit your site, find content gaps, and create briefs your team can ship.