In brief Share of voice in AI answers is a sample statistic, not a census of the internet. It only means something when the prompt panel, the coding rules, and the unit of observation are written down. Combining mention, citation, and recommendation into one number usually hides the decision you care about.

“Share of voice” is a borrowed phrase. In advertising it once meant a brand’s portion of paid impressions. In media analysis it often means a brand’s portion of coverage in a defined set of outlets. In AI answers it should mean something stricter: a brand’s share of a coded event inside a defined prompt panel, on a named engine, during a dated collection window.

If any of those pieces is missing, the percentage is a slogan.

This brief is about how to keep the measure from collapsing. It does not report a Northline league table. It describes the instrument.

Start with the unit, not the dashboard

Decide what event you are counting before you decide how to chart it.

Event Question the coder asks Typical use
Mention Does this answer name the brand or product? Awareness in the category
Citation Does the interface attach this brand’s URL (or a page clearly about it) as a source? Evidence and discoverability
Recommendation Does the answer tell the user to consider, prefer, or start with it? Commercial outcomes

A brand can win mentions and lose recommendations. A documentation site can win citations and never be named as a vendor. Those are different stories. Averaging them into “AI visibility” makes the story cheaper and less true.

If a stakeholder wants a single index, build it as a labeled composite and show the parts. Do not let the composite become the only number in the room.

The prompt panel is the sample frame

AI answers are not a stable corpus like last month’s news archive. They are generated at request time, for a prompt, with a model version, tools, and location context you may not fully see. Your “population” is therefore the set of questions you chose to ask.

That choice is the study.

A panel that only contains “best [category] tools” questions will over-represent recommendation language and comparison blog sources. A panel that only contains “what is [brand]” questions will over-represent owned pages. A panel written by the sales team will over-represent the way the company wishes buyers spoke.

Write the panel as if it were a survey instrument:

  • State the buyer jobs and question types you intend to cover.
  • Include unbranded category questions and a smaller set of branded questions, and keep those sets separate in the results.
  • Version the panel. If you add or drop prompts, say so, because the time series breaks.
  • Record engine, interface, date, and any settings you can see (web browsing on or off, location, logged-in state).

Repeating the same 12 prompts every quarter is more informative than inventing 80 new ones and calling the movement “the market.”

Coding has to survive a second reader

Share of voice is only as good as the codebook. Ambiguous names, product families, and “the category leader” language will otherwise be scored by vibe.

A workable codebook answers at least these questions:

  • What strings count as the brand (legal name, product line, common shorthand)?
  • Do parent and subsidiary names count as the same entity?
  • Is a mention in a refusal, a disclaimer, or a “we couldn’t find” clause still a mention?
  • Does a citation to a third-party review that discusses the brand count as the brand’s citation, or as the publisher’s?
  • What words count as recommendation (“consider,” “popular option,” “if you need X”) and what words are merely descriptive?

Two people should be able to apply the rules to the same answer and agree most of the time. When they do not, the disagreement is data. It tells you the definition is soft.

We do not treat a model as a coder of record. Models can pre-label. People adjudicate the hard cases and own the published numbers.

Platforms are not interchangeable slices of one pie

ChatGPT, Perplexity, Gemini, and Google AI Overviews do not produce the same object. Some answers look like essays with optional footnotes. Some look like sourced research notes. Some sit inside a search results page and compete with links, ads, and sitelinks.

You can still use the same event definitions across engines. You should not add their counts together and call the sum “AI share of voice” unless you have a reason that survives a sentence. A citation in an AI Overview is not the same exposure as a named recommendation in a chat thread.

Report each engine on its own line. If you need a cross-engine view, show a small-multiples chart, not a blended percentage that no one can audit.

What not to do with the number

Do not treat a one-day collection as a run rate. Outputs move with model updates, retrieval freshness, and even the wording of a follow-up.

Do not convert share of voice into revenue without a separate study of click-through, assisted search, or sales. The answer engine is not obligated to send traffic, and many answers satisfy the question on the page.

Do not hide zeros. A brand that never appears in an unbranded panel is a finding. Padding the panel with branded prompts to “give the brand a chance” changes the question you are asking.

Do not invent precision. If you ran 40 prompts, do not talk as if you measured the entire category to a tenth of a point. Publish the panel size beside the rate.

A minimum report

A share-of-voice note that we would stand behind includes:

  1. The event definition (mention, citation, recommendation, or a named composite).
  2. The panel (count, jobs covered, branded vs unbranded split, version).
  3. The engines and dates.
  4. The codebook and any inter-rater check.
  5. The rates by brand and by engine, with zeros left visible.
  6. A short list of what the study did not measure (traffic, conversions, paid placement, offline conversation).

That package is slower to produce than a screenshot. It is also the difference between research and a mood.