Most companies have never actually looked. They assume the answer is no, or assume it does not matter yet, and in the meantime a competitor is being named in the answer their buyers are reading.
Before any work to change that is worth commissioning, it is worth knowing where you stand. That turns out to be a measurement problem with several traps in it, and most informal attempts fall into at least one. How to get cited by ChatGPT covers what actually influences citation. This is about finding out what is happening now, and keeping track of it honestly.
Why does one query tell you almost nothing?
Ask an assistant the same question twice and you can get different companies, in a different order, citing different sources. These systems sample rather than look up a fixed ranking, retrieved results shift, and model updates change behaviour without notice.
So a single check is an anecdote, not a finding. The thing that can actually be measured is a pattern across a fixed set of questions, sampled repeatedly over time. Any individual answer, flattering or not, should be treated as noise until it repeats.
What quietly biases the result?
This is the trap that invalidates most do-it-yourself audits, and it is worth getting right before you record a single result.
If you run these questions in your own logged-in account, the assistant may personalise the answer. ChatGPT draws on saved memories and can reference your past conversations. If you have spent the last month discussing your own company in that account, being named in the answer tells you nothing about what a stranger would see.
Two practical steps:
- Use a Temporary Chat. It starts non-personalised, does not write new memories, and does not join your history.
- Check what informed the reply. ChatGPT can show whether a response drew on memory, past chats, custom instructions or files. If it used any of them, discard that result and run it clean.
Then fix and record the conditions: which assistant, whether web browsing was on, roughly where you are, and what language you asked in. Answers legitimately differ across all four, and a comparison across months is worthless if those were drifting underneath it.
Which questions should you actually test?
The ones buyers genuinely type, phrased the way they phrase them. Three shapes are worth covering, because they catch different stages of a decision:
- Category questions — "best [service] for [type of company]".
- Comparison and alternative questions — "alternatives to [competitor]", "[competitor] vs [competitor]".
- Problem-shaped questions — "how do I…", asked before the buyer has settled on a category at all. These are the ones brands most often miss, because they contain no product language to optimise against.
Fifteen to thirty questions is plenty. Fewer and you are reading noise. Many more and you will not re-run it, which defeats the point.
One question to skip: your own brand name. "What is [company]" almost always returns something serviceable, and it tells you nothing about whether you would be found by someone who has never heard of you. Discovery is the thing being measured, not recall.
What should you record?
Three outcomes get lumped together as "did we show up", and they are different problems with different fixes:
- Absent — competitors are named, you are not.
- Mentioned — named in the text, with nothing linked back to you.
- Cited — named and carried as a linked source.
Alongside each one, log the date, the assistant, whether browsing was on, which competitors appeared, and — most importantly — which sources the answer cited. A spreadsheet is entirely sufficient. The value is not in any single row; it is in comparing identical questions across months.
The most useful output is not your own score
It is the list of sources.
Across a full question set, the cited sources tend to converge on a fairly small group of publications the model already treats as credible for your category. That list is evidence, not assumption, about where independent coverage would plausibly change what the assistant says.
Most brands find that more actionable than their own mention rate, because it converts a vague goal — "we want to show up in AI answers" — into a specific, named set of outlets worth earning coverage in. It also tends to be shorter, and less predictable, than the list a team would have guessed.
How often should you re-run it?
Monthly is a sensible default. Quarterly is the floor. Weekly is too often: ordinary run-to-run variance will look like movement, and you will start reacting to noise.
Re-run the same questions, unchanged. Rewriting them each round feels like an improvement and quietly destroys comparability, which is the only thing this exercise actually produces.
What this cannot tell you
Being straight about the limits matters, because plenty of tools sell past them:
- It is a sample, not a ranking. There is no public position and no share-of-voice figure to report, whatever a dashboard implies.
- There is no volume data. You cannot see how many people asked, so you cannot weight one question against another by demand.
- It is assistant-specific. Results on ChatGPT say little about Google's AI surfaces, which lean far more heavily on conventional search rankings — see why ranking on Google doesn't mean ChatGPT cites you and how to appear in Google AI Overviews. If Perplexity matters to your buyers, it behaves differently again.
- Nobody controls the output. There is no submission form and no paid inclusion. Anyone offering guaranteed placement in an AI answer is describing something that does not exist.
What do you do about a gap?
Gaps come in two kinds, and telling them apart is most of the value of running the audit at all.
A content gap means the question simply is not answered properly anywhere you publish. That is yours to fix, and it is the faster of the two.
A corroboration gap means the answer exists on your site, but nothing independent supports it. Any company can write anything about itself, so a claim that appears nowhere else gives an assistant little reason to repeat it. This is slower to close, and it is usually the one deciding whether you get named.
Most brands have the second problem and try to fix it with the first. That is why so much AI-visibility work produces a tidier website and no change in the answers.
If the audit points at corroboration rather than content, that is what our AI visibility service is built to address: the structural work and the earned coverage together, rather than the checklist half on its own.

