When Google shows an AI Overview, it answers the question directly on the results page and cites a small set of sources beside the answer. For a brand, the stakes are simple. The cited sources are part of the answer, and everything else appears below it. The practical question for a marketing leader is how the system decides which brands to cite.
Google does not publish the criteria. The pipeline that produces an AI Overview is partly documented and partly observable, and it rewards specific things a brand can control. Understanding the mechanism is more useful than chasing any single tactic, because it explains why the tactics work.
How an AI Overview is assembled
An AI Overview is the output of a pipeline, and each step of that pipeline is a filter a brand can pass or fail.
Query fan-out. Google has said publicly that its generative search features run multiple related searches behind a single question, a process it calls query fan-out. A question like “should I use a mortgage broker or go direct to a bank” may spawn sub-queries about broker commissions, lender panels, approval odds and regulation. The overview is assembled from the results of all of them, including the ones the user never saw.
Retrieval from the index. Candidate material comes from the same index that powers ordinary search. A page that is not crawlable, not indexed, or excluded from snippets cannot be used. This is why technical SEO remains the entry ticket. The generative layer only retrieves what the index already holds.
Passage selection. The system works at the level of the passage. From the retrieved documents it pulls sections that appear to answer a specific sub-question: a clean definition, a direct comparison, a stated fact. A page is not cited as a whole; a passage from it is.
Synthesis. A language model composes the answer, grounded in the retrieved passages. Grounding is the operative constraint. The model is steered towards statements its retrieved sources support, because unsupported generation is the failure Google is most visibly trying to avoid.
Citation selection. Links are then attached to parts of the generated answer. The sources cited are the ones whose passages the answer drew on, so the citation list behaves like a bibliography.
Each step filters the field. A brand can fail at retrieval, fail at passage selection, or pass both and still lose at citation because a competitor’s passage supported the claim more cleanly.
The signals that correlate with citation
Nobody outside Google can list the weightings, and anyone claiming otherwise is selling something. The sources that keep being cited do share observable traits, and the same traits show up across ChatGPT, Perplexity and Copilot, which face the same grounding problem.
Passage-level answers. Content that answers one question per section, in complete declarative sentences, gives the system something to lift. A definition stated plainly in two sentences can be quoted directly. Narrative that spreads the same information across three paragraphs gives the system nothing to take.
Entity clarity. The systems reason over entities: who the brand is, what it does, and where it operates. Schema markup, consistent naming and unambiguous organisational facts make the brand something the model recognises. Ambiguity is expensive, because a model unsure which entity it is describing tends to describe a different one.
Corroboration across sources. A generated answer repeats claims at scale, so the safest claim to repeat is one that multiple independent sources agree on. Facts that appear only on a brand’s own site are weaker candidates than facts echoed by directories, press coverage and third-party commentary. Corroboration is how the pipeline approximates trust.
Freshness. Retrieval favours current material, particularly where the question implies recency: pricing, regulation, and anything with a year in it. A page that was authoritative when written and untouched since will keep its ranking longer than it keeps its citations.
None of this is published as a recipe. It follows from how a grounded generation system has to behave: retrieve candidates, then prefer the passages that are extractable, corroborated and current.
Why ranking first does not guarantee citation
The most common misreading of AI search is that citation is a reward for rank. Four mechanical reasons say otherwise.
Ranking is a page-level judgement about relevance to the visible query. Citation is a passage-level judgement about usefulness to a specific sub-question the user never typed. A page can be the best overall result and still contain no passage that cleanly supports any sentence of the generated answer.
Fan-out widens the field. The overview draws on queries adjacent to the visible one, so sources that rank for the sub-queries enter the candidate pool even when they do not rank for the headline term. A brand watching only its primary keyword cannot see most of the contest it is in.
Extractability beats position. When two candidate passages support the same claim, the plainly stated one is easier for the system to use and attribute. A top-ranked page with its answer buried will lose to a mid-ranked page that states the answer in a sentence.
Corroboration filters late. A claim the model cannot see supported elsewhere is a risk to repeat, however well the page hosting it ranks.
What brands can control
Nobody can control the model. What it retrieves is another matter, and that comes down to a defined body of work.
Structure content so each commercially important question is answered in its own section, in sentences that survive being quoted out of context. State the facts directly: what the service is, who it is for, where it operates. Mark up the organisation, its people and its services with schema, and use the same names and descriptions everywhere. Reconcile the facts across the site, directories, profiles and press, because every contradiction gives the system a reason to hedge or to cite someone else. Earn third-party corroboration for the claims that matter most, since a claim echoed independently is safer to repeat. Keep the pages that answer buying questions current, and show the date.
Then measure it directly. Run the buying questions through the engines on a schedule and record who gets cited, how the brand is described, and which competitors appear. Rankings are no longer a proxy for this, so the reliable read on citation is to ask the engines.
A practical checklist
- List the questions buyers ask before choosing a provider in your category, including the ones they would never type as a keyword.
- Check whether each question has a page, and whether that page answers it in a liftable passage near the top.
- Validate schema for the organisation, services and FAQs, and fix naming inconsistencies across the web.
- Compare how your site, your directories and your press describe the business; reconcile the conflicts.
- Identify your most commercially important claims, and find or build independent corroboration for them.
- Date and refresh the pages that answer buying questions, so retrieval sees them as current.
- Query AI Overviews, ChatGPT, Perplexity and Copilot with your buying questions monthly, and record the citations.
Done once, this is an audit. Repeated on a schedule, it becomes AI search optimisation: the ongoing work of making sure your brand is part of the answer when the engines assemble one in your category.