How to read an AI visibility study
Five questions that separate a measurement from a marketing asset. Use them on us too.
methodological survey of 45 studies, July 2026 · July 8, 2026 · 4 min read
You are going to be shown studies. Every vendor in this market publishes them, we publish them, and most of them are produced by companies with a commercial interest in the conclusion. That is not disqualifying — some of the best data in this field comes from vendors, because they are the only ones with the data — but it means you need a way to read them.
In July 2026 a methodological survey reviewed 45 studies of generative engine optimisation published between November 2023 and July 2026. Its criticisms of the literature make an unusually good checklist, and we have used it on our own work.
1. Does it presuppose retrieval?
The most common flaw. A study places five documents in a model’s context, changes one, and measures which gets cited. That answers “given that you were found, were you chosen?” — and skips the harder question, which is whether you were found at all.
The survey calls this fixed-context bias, and it matters because the two can point in opposite directions. One benchmark of 171,003 documents found that body-only optimisation reduced top-20 presence by about 9% and final citation by 6% — tactics that won the citation coin-flip lost the earlier round. Winning the flip is worthless if you fell out of the candidate set.
Ask: was the document retrieved, or was it handed to the model?
2. Who is judging?
A recurring pattern: the same family of models generates the test content, rewrites it, and then evaluates which version is better. The survey calls this LLM-judge circularity. It is not fraudulent; it is just a closed loop, and it will reliably find that content optimised for models is preferred by models.
The strongest study in this field — 252,000 trials, the price-disclosure result — has exactly this property, and its authors say so. They hand-reviewed 300 scenarios as a check. We cite that study repeatedly and it has this flaw, which is why we label it as a preprint every time.
Ask: who generated the corpus, and who scored it?
3. What is the denominator?
“Share of citations” is often computed only across answers that contained citations. But most answers don’t. Similarweb found that only about 6.8% of US ChatGPT prompts produced an answer containing any citation in May 2026 — up sharply from 1.6% a year earlier, and still small.
So a headline saying a brand holds 30% share of citations may describe 30% of a very small number. That is a selection bias, and it is invisible unless you ask.
Ask: 30% of what, and how many answers had no citations at all?
4. Is anyone else optimising?
Almost every study tests one actor changing one thing while the rest of the world stands still. Real markets don’t work that way. C-SEO Bench — peer reviewed at NeurIPS 2025, which most work in this field is not — found that most conversational-SEO methods “are not only largely ineffective but also frequently have a negative impact on document ranking”, and that effectiveness declined as more competitors adopted the same technique. The authors describe the space as congested and zero-sum.
Which means a tactic that worked in a study run in early 2025 may not work now, for the boring reason that everybody read the same study.
Ask: what happens when your competitors do this too?
5. Is it measuring visibility, or a proxy for visibility?
The survey’s bottom line is worth quoting whole:
“No reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior.”
Citation counts, mention rates, share of voice — these are proxies. The thing anyone actually wants is customers. The chain from one to the other has not been demonstrated by anyone, in this field, yet.
Ask: what is the outcome being measured, and how far is it from money?
Now apply it to us.
Our free report asks twelve questions across three models, twice each. Here is how it scores against its own checklist.
Retrieval: we measure live answers to real questions, so retrieval is included rather than assumed. Good.
Judging: a model extracts the mentions and positions from each answer, but it never computes a score — the arithmetic is done in code, from stored data. That avoids the worst of the circularity, though the extraction step is still a model reading a model.
Denominator: we report the number of answers that named nobody at all, alongside the ones that named someone. That number is often the most interesting one on the page.
Competition: we can’t control for it. Nobody can. If everyone in your category does this work, the advantage compresses. We would rather say so than sell you permanence.
Proxy distance: we measure whether a machine names you. We do not measure whether that produced a customer, and we will not claim to. That is the honest boundary of what this service can prove, and any vendor telling you otherwise is describing a causal chain the literature has not established.
One last note on grading. Most of the strongest results in this field — including several we rely on — are arXiv preprints that have not been peer reviewed. We label them every time. When someone shows you a study without saying which category it falls into, that omission is itself information.