
AI share of voice is your proportion of all brand mentions across a fixed set of tracked prompts, measured against a defined competitor set. It answers the question executives actually ask, which is not whether you are visible but whether you are more visible than the companies you lose deals to.
Citation frequency tells you how often you appear. Share of voice tells you how often you appear relative to everyone else named in the same answers, which is the number that tracks competitive position.
The distinction matters because both can move independently:
Neither number is superior. Frequency answers whether the work is producing exposure; share answers whether that exposure is winning ground. A program reporting only one of them will eventually be surprised by the other.
Share of voice is the metric that survives a challenge from finance. It is comparative, it uses a competitor set the business already recognizes, and it cannot be inflated by adding prompts you already win.
That last property is worth spelling out. Citation frequency can be improved by adding easy prompts to the set, which makes it a metric anyone can game without meaning to. Share of voice cannot, because adding a prompt you win also adds every competitor named in that same answer to the denominator. It is self-correcting in a way that most marketing metrics are not.
Build the set from the answers, not from your internal view of the market. This is the step that changes people’s minds.
Two findings recur. Brands routinely discover an unfamiliar competitor appearing consistently in their category, and they routinely discover that a rival they benchmark against obsessively is invisible in AI answers. Both change how budget gets allocated, and both come from the same 90 minutes of testing.
Keep the set fixed for the reporting period. Adding a competitor mid-quarter changes the denominator and breaks comparability, which means an apparent drop in your share may only reflect a change in who you are counting. Review the set quarterly, document any change, and treat the period after a change as a new baseline rather than a continuation.
The arithmetic is simple. The discipline is in keeping the inputs constant.
| Input | Definition |
| Prompt set | Fixed wording, 30 to 50 prompts, frozen between periods |
| Runs | Three or more per prompt per platform |
| Competitor set | Fixed for the reporting period |
| Mention | Brand named anywhere in the answer |
| Share | Your mentions divided by total mentions of all tracked brands |
Freezing these inputs is what makes the metric usable. A share figure calculated on a different prompt set, a different number of runs, or a different competitor list is a different metric that happens to share a name, and comparing the two produces conclusions that are confidently wrong.
Calculate it three ways, because each answers a different question:
The platform split is where most of the useful detail sits. A brand can hold 20% share on ChatGPT and zero on Gemini, and an overall figure of 11% describes neither situation.
One methodological decision needs making before the first calculation: whether a brand mentioned twice in a single answer counts once or twice. Either convention works as long as it never changes. Counting once per answer is simpler and less sensitive to answer length, which makes it the safer default for a metric that has to stay comparable across platforms with very different verbosity.
Not every change means something. Interpret against three references.
| Observation | Likely explanation |
| Small movement inside normal run variance | Noise, not signal |
| Whole set moves at once, all brands affected | Platform change, not performance |
| Your share rises on one prompt type only | Work in that area is landing |
| A single competitor gains across everything | They shipped something, or earned significant coverage |
| Your share falls while frequency holds | New entrants diluting the field |
Record platform changes in a run log. Default models get swapped and retrieval behavior shifts, and when that happens the useful record is the note explaining what changed rather than the movement itself. Comparing SEO and AI search in 2026 makes the same point structurally: the ground moves under the measurement more often than it did in traditional search.
Establish what counts as meaningful movement before you need to interpret one. If your run-to-run variance on a stable prompt set is five percentage points, then a four-point quarterly change is noise and should be reported as flat. Teams that skip this step end up narrating random variation to executives, which costs credibility precisely when the genuine movement arrives.
The aggregate number gets attention. The prompt-level breakdown produces the work.
For every prompt where a competitor appears and you do not, classify why:
The distribution is the finding. A set dominated by citability gaps means your content is adequate and badly structured, which is the cheapest position to be in. A set dominated by corroboration gaps means no amount of additional content will change the outcome, which is the most expensive thing to learn late. Building an AI content strategy from that distribution allocates budget better than building one from a topic list.
Also record which domains were cited as evidence on each lost prompt. That column identifies the specific sources supplying your category’s answers, which converts a vague instruction to build authority into a named list of five or six places to work on. It doubles as competitive intelligence, since the evidence behind a competitor’s recommendation shows where their standing actually comes from.
Share is a ratio, so it falls when the denominator grows. Three situations produce this.
Always report share and frequency together. Either number alone can tell a misleading story, and the pair is unambiguous: rising frequency with falling share is a competitive problem, falling both is an execution problem.
The inverse case is worth naming too. Share can rise because competitors deteriorated rather than because you improved, which looks identical in a share-only report. Tracking each competitor’s absolute frequency alongside your own separates the two, and it prevents a team taking credit for a rival’s neglect and then being surprised when that rival resumes investing.
Prioritize by winnability rather than by volume, since there is no volume data to prioritize by.
| Signal | Interpretation |
| Current answer is vague or generic | Winnable, the incumbent answer is weak |
| Same three brands every run | Entrenched, expensive to displace |
| Competitors appear inconsistently across runs | Unsettled answer, good target |
| Cited evidence comes from sources you can reach | Winnable through corroboration work |
| Cited evidence comes from sources you cannot reach | Deprioritize for now |
| You appear but are described inaccurately | Entity work, not content work |
Work the winnable prompts first even where they are commercially smaller. Early movement funds the slower corroboration work politically as well as financially, and a program that shows nothing for two quarters rarely survives to see the gains it was actually built for.
Category concentration sets the realistic ceiling. Semrush found the three most visible brands hold 82.9% of category visibility in News and Media and 76.9% in Consumer Electronics, against 42.2% in Industrial and 41.4% in Finance. In concentrated categories the winnable prompts are narrow and specific; in distributed ones, consistent execution can move share within two quarters.
Set the target against that ceiling rather than against a round number. A goal of 25% share in a category where three incumbents hold 77% is not ambitious, it is arithmetically implausible, and committing to it guarantees a program is judged a failure while performing normally. A defensible target names both a share figure and the specific prompt types it applies to.
A share of voice report needs to survive someone asking how the number was produced.
Include, every time:
Exclude:
Semrush found that 45% of marketing leaders cannot accurately measure brand visibility inside AI-generated answers, and only 9% have tools covering all relevant metrics. A transparent method beats a sophisticated-looking number in that environment, because the method is what makes the number defensible.
The presentation that lands best pairs the share figure with two or three named prompts. A slide stating that an assistant recommends three competitors and never mentions you when asked a specific buying question does more work than any percentage, because it converts an abstract metric into a sentence a sales leader recognizes from their own pipeline.
| Activity | Cadence |
| Full prompt set run | Monthly |
| Share of voice reported | Quarterly |
| Competitor set reviewed | Quarterly |
| Prompt set refreshed with new language | Quarterly |
| Full rebuild and new baseline | Annually, or after repositioning |
Monthly running with quarterly reporting is the right split. Monthly data smooths run variance; quarterly reporting matches the pace at which the underlying work actually produces movement, and it avoids the trap of explaining noise to an executive audience every four weeks.
Two events justify an unscheduled measurement. A significant platform change, such as a new default model, can move an entire prompt set at once and is worth capturing while the before-and-after is clean. And a competitor launch or major funding announcement frequently precedes a coverage surge, which shows up in share of voice weeks before it shows up anywhere else you would notice.
Build the first report with the intention of repeating it exactly. The format matters less than its repeatability, since the value of this metric accrues over quarters rather than within one, and a report redesigned each period loses the comparison that justified building it.
Share of voice is the number to put in front of leadership. It is comparative, method-transparent, and directly tied to the competitors the business already worries about, which is why understanding why answer engines prioritize certain brands matters more than any single score. A generative engine optimization program should produce this baseline before proposing work, and the composition of the competitor set is worth agreeing together when you get in touch.
How is AI share of voice different from citation frequency?
Frequency counts how often you appear across your tracked prompts. Share of voice divides your mentions by the total mentions of all tracked brands in the same answers. Frequency measures exposure; share measures competitive position, and the two can move in opposite directions.
How do I choose which competitors to track?
Derive the set from the answers rather than from internal assumptions. Run your prompt set, log every brand named, rank by frequency, and take the top five to eight. Brands you consider rivals that never appear are not competing for the same visibility, which is useful information in itself.
Why did my share drop when my mentions increased?
Because share is a ratio. More brands entering the category, or a platform citing more sources per answer, both increase the denominator. This is why share and absolute frequency should always be reported together, since either alone can describe the same quarter as a success or a failure.
Can I compare my share of voice to published industry benchmarks?
Not reliably. Published studies measure different things, with some counting the share of answers citing a domain and others counting each domain’s share of all citations, which produces figures differing by an order of magnitude. Compare against your own prior periods instead.
How long before share of voice moves?
Citability improvements can show within a quarter, since they affect content you already control. Corroboration-driven gains typically take two quarters or more, because they depend on third-party sources publishing on their own schedules rather than yours.