In brief Citation is not the same as mention. A brand can appear in the prose of an answer without being used as a source. Researchers can study visible patterns in outputs. They cannot reconstruct a complete ranking formula from those outputs alone.

Answer engines do not publish a complete account of how they choose the pages they cite. That does not make citation selection a mystery you have to guess at. It means researchers should separate what can be observed from what is being inferred.

This note covers four things that can be studied in public outputs: the difference between mention and citation, the kinds of sources that tend to appear, the role of the question, and the limits of any single snapshot.

Mention, citation, and recommendation are different events

A model can name a company in a sentence, point to a URL as support, or tell the reader to use a specific product. Those are different events.

  • A mention is a name appearing in the answer text.
  • A citation is a source the interface treats as evidence, usually with a link, footnote, or source card.
  • A recommendation is an instruction or ranking (“start with,” “prefer,” “best for”) that goes beyond naming.

A visibility program that folds all three into one score will hide the thing that matters. A legal publisher may want citations. A consumer brand may care more about recommendation. A challenger brand may only be trying to appear at all.

When we code answers at Northline, we keep the three events on separate fields. If a later chart combines them, that combination is a derived view, not the raw observation.

What source selection looks like from the outside

Public products change, and vendors do not document full retrieval stacks. Still, a few regularities show up often enough that they are useful working hypotheses, not laws.

Consensus and corroboration. Questions with a settled factual core tend to pull sources that agree with one another. Encyclopedia pages, official documentation, and widely mirrored explainers appear because they are easy to retrieve and hard to contradict. That is not proof that “authority” is a single score. It is a reminder that lonely claims are weaker evidence for a model that is trying to avoid being obviously wrong.

Primary and official sources. Policy, pricing, product capability, and regulated topics often surface issuer sites: a government page, a company help center, a standard, a filing. If your entity is not clearly identified on a page that an engine can retrieve, you are asking the model to invent a source or skip you.

Recency, but only for some questions. Time-sensitive prompts (“what changed this week,” “current pricing,” “latest outage”) reward freshly updated pages. Evergreen definitional prompts do not. Treating every query as a recency contest wastes effort and produces noisy audits.

Page shape. Answers need extractable passages. Titles that match the question, short definitional openings, tables, and FAQ-style sections are easier to lift than a brochure that never states the fact. This is not a claim that structured data “wins AI search.” It is a narrower point: if the passage is hard to find, it is hard to cite.

Identity. Engines cite pages. They also have to decide which real-world entity a string refers to. Ambiguous names, reused product labels, and thin “about” pages make that matching harder. Citation problems are often entity problems.

None of these observations replaces a retrieval paper from the vendor. They are coding categories. If a study cannot point to the output that justified a label, the label should not be in the report.

The question is part of the ranking

Citation mix changes when the prompt changes. A “what is” question, a comparison, a local service question, and a “best tools for” question are different tasks. They invite different source classes.

That is why a single hero prompt is a poor instrument. It is also why screenshots of one ChatGPT session are not a market study. If you want to talk about how an engine chooses sources, you have to say which job the user was trying to do.

A practical way to keep this honest:

  1. Write the user job first (learn, compare, buy, troubleshoot, verify).
  2. Write the prompt second.
  3. Code citations against that job, not against a generic idea of “authority.”

A government domain cited on a troubleshooting prompt may be a miss, even if it looks prestigious. A forum thread cited on a “how do people actually do this” prompt may be a fit.

What you cannot see

Most production stacks combine some mix of web retrieval, site indexes, licensed data, safety filters, and model memory. The interface does not tell you which layer produced a sentence. A citation can be decorative, supporting, or loosely related. Some answers cite nothing and still sound certain.

So a citation study should report absences as well as presences. “No sources shown” is a finding. “Sources shown but unused in the prose” is a finding. “The same URL cited for unrelated claims” is a finding. Pretending that every footnote is a clean attribution trail is not research.

It is also a mistake to treat yesterday’s pattern as a durable ranking factor. Interfaces, retrieval partners, and models move. The durable object is the method: a dated prompt panel, a documented codebook, and a clear statement of what was not measured.

A working checklist

If you are reading an engine’s citations rather than chasing a secret algorithm, these questions are enough to start:

  • Was the brand mentioned, cited, recommended, or some mix?
  • What user job was the prompt actually asking?
  • Which source classes appeared (official, reference, news, vendor, community)?
  • Did the cited passage support the sentence it sat beside?
  • What would have to be true on the page for this entity to be citable next time?

Those questions will not give you a published ranking function. They will keep a visibility program attached to evidence.