Most executives can pull up their search rankings in seconds. Very few can say whether their brand appears when a buyer asks ChatGPT, Perplexity or Google’s AI Overviews who to use. You can measure that gap yourself. It takes a defined query set, a spreadsheet and a few hours a month, with no tool subscription and no agency.
What follows is the full method, laid out so you can run it yourself. It is the same logic a professional audit uses. An AI search visibility audit measures three things: whether a brand is cited in AI-generated answers, how it is described when it appears, and which competitors are cited when it is not. A fourth question follows all three: why. Everything below serves those four questions.
Choose the queries that matter
The audit is only as good as its query set, and the most common mistake is auditing the wrong questions. Keyword rankings are the wrong starting point. Build the set from the decisions buyers make on the way to choosing a provider in your category, because those are the questions they now put to AI engines.
Use four intent types:
- Category questions. How the buyer frames the problem before they know the solution: “how do I get a home loan approved with irregular income”, “what does a marketing audit involve”. These show whether you are visible early, while preferences are forming.
- Comparison and shortlist questions. “Best mortgage broker in Brisbane”, “top providers of X”, “A vs B”. These are the money queries. An engine names two or three brands here and effectively writes the shortlist.
- Evaluation questions. What buyers ask once a shortlist exists: pricing questions, “is X worth it”, “what should I look for in a provider of Y”. Being present here shapes the final decision.
- Brand questions. Your own name, plainly: “who is [brand]”, “is [brand] reputable”, “[brand] reviews”. You will almost always appear for these, so the audit question is whether the description is accurate and current.
Fifteen to thirty queries is the practical range for a self-run audit. Fewer than fifteen and the sample is too thin to read patterns from. More than thirty and the monthly time cost gets high enough that the discipline collapses. Weight the set towards comparison and evaluation queries, because that is where citation converts to revenue, but keep at least a few from each intent type. Write the queries the way a buyer would phrase them, as full questions in natural language. Then freeze the set. The value of the audit comes from running the same queries over time.
Run them across the four engines
Run every query through each of the four engines that currently matter for Australian buyers: Google AI Overviews, ChatGPT, Perplexity and Microsoft Copilot. Use the same query set across all four so the results are directly comparable.
A few rules keep the data honest:
- Reduce personalisation where you can. Use a private or logged-out browser session for Google, and be aware that ChatGPT and Copilot results still vary with account history and location. You cannot remove personalisation entirely, so note it as a limitation of the run.
- Use the query verbatim. Do not rephrase when an engine misunderstands the question. A buyer would not rephrase around your brand’s blind spot, and the misunderstanding is itself a finding.
- Capture evidence as you go. Screenshot or export each answer. When a citation appears or disappears next month, you will want to see exactly what changed.
- Record non-answers too. If a query produces no AI Overview, or an engine declines to name providers, that is a finding. It tells you where the generated-answer contest has not started in your category.
Expect the full pass to take a working half-day the first time, and less once the routine is set.
Record it in a fixed template
You do not need software for this. A single spreadsheet does the job, one row per query per engine per run, provided the columns are fixed before you start. Record, for every row: the date of the run, the engine, the exact query, whether an AI-generated answer appeared at all, whether your brand was cited or named in it, how prominently (named in the answer text, cited as a linked source, or both), a verbatim note of how the brand was described, every competitor named or cited, and the sources the citations pointed to. That last column matters because the engines often cite publishers and directories, and those intermediary sources are where visibility is frequently won or lost.
Two of those columns deserve discipline. Copy the description note word for word. Drift in how an engine describes you is one of the most actionable findings an audit produces, whether that is an old service line, a wrong location or a stale positioning. The competitor column should include everyone named, including rivals you did not expect, because the engines routinely surface competitors a brand has never benchmarked against.
Score it and set the baseline
Raw rows become useful when you compress them into a small set of numbers you can track. Three numbers are enough.
Citation share. The percentage of query-engine combinations where your brand appears in the generated answer. It is the AI search equivalent of a rankings report: one number, tracked over time, that tells an executive whether visibility is improving. Calculate it overall and per engine, because the engines behave differently. A brand can be strong in Perplexity and invisible in AI Overviews.
Description accuracy. Of the answers where you do appear, the share where the description is accurate and current. A three-level grade is enough: accurate, partially accurate, wrong. An engine that cites you with outdated or incorrect information can do more damage than not appearing at all.
Competitor citation frequency. A count of how often each competitor appears across the set, ranked. The top of that list is the competitive map for AI search in your category, and most organisations have never seen it.
The first run is the baseline. It will probably be uncomfortable to read. Every later run is measured against it, so AI visibility becomes a tracked metric.
Repeat it on a cadence
One run is a snapshot. The value is in the series. Monthly suits most organisations, because it catches movement without taking more time than a team will keep giving it. Keep the query set frozen between runs. When you do add queries, add them alongside the original set so the baseline remains comparable.
Cadence matters for a second reason. AI answers are non-deterministic. The same query can produce different citations on different days, so visibility is properly measured as a rate over repeated runs. A citation that appears in one run and vanishes in the next usually shows an unstable position, and only a series will reveal that.
Read the results as a diagnosis
The numbers tell you where you stand. Working out why takes a model of how the engines choose. Four factors do most of the explanatory work: content structure, entity clarity, authority signals, and consistency of facts.
If competitors are cited for questions you have never answered in a dedicated, liftable passage, the gap is content structure. The engines extract passages, so they cannot cite an answer you have not written. If the engines describe you vaguely, confuse you with another organisation, or hedge on basic facts, the gap is entity clarity: the schema, naming and organisational facts that let a model treat you as a known thing. If your content answers the question well but the citations keep going to publishers and directories, the gap is authority. The engines are choosing sources they already trust, and your version of the facts lacks third-party corroboration. If your description varies engine to engine, check what the web says about you. Conflicting facts across your site, directories and coverage give a model reason to hedge, or to cite someone whose story holds together.
Why models select the sources they do is covered in more depth in what makes a model cite one brand over another. For the audit itself, the four factors are enough. Every gap the spreadsheet surfaces will trace back to at least one of them, and that is how a measurement exercise turns into a work program.
The honest limits of doing it yourself
This method is real and the results are usable. It also has limits worth stating plainly, because they are the difference between a self-run check and a professional audit.
Query sampling. Fifteen to thirty queries is a sample of a much larger space. The engines fan a single question out into many related searches, and buyers phrase the same intent dozens of ways. A small sample can miss the phrasings where the real contest is happening, and choosing a genuinely representative set is harder than it looks from inside the brand.
Personalisation and variance. Logged-out sessions reduce personalisation without removing it, and they do nothing about run-to-run variance. Separating signal from noise across four engines takes statistical care, or enough repeated runs for the patterns to settle. Both cost time.
Time cost and interpretation. The recording is mechanical. The reading takes judgement: tracing a citation gap to its cause, ranking the fixes by commercial impact, and deciding which gaps are worth closing at all. This is where self-run audits most often stall, with a spreadsheet full of findings and no sequenced plan.
None of that is a reason to skip the method. Run it. A brand that measures its own AI visibility monthly is already ahead of most of its category. If the baseline is confronting, or the findings need to carry weight with a board, the fixed-scope AI search visibility audit is the senior-led version of the same method. It runs the same questions with a defensible query set and documented evidence, and it ends in a competitive read and a roadmap ranked by commercial impact.