How to measure this properly
Everything else on this site depends on one thing being possible: that you can find out what the engines actually say about you, repeatably, and hand the same method to someone else. This is that method.
It is also the standard the leaderboard scores agencies against, so a company doing all of this scores well and a company doing none of it scores zero. Nothing here requires a licensed product. A spreadsheet and an afternoon a month will do it.
The seven parts
- A fixed prompt set
- Written down before the work starts, agreed in writing, and unchanged between reports. This is the denominator of every number that follows. See prompt set.
- Named engines
- Named individually, with the model version where the engine exposes one. “AI” is not an engine and a report covering “AI platforms” is covering something you cannot audit.
- A stated location, logged out
- From the market you actually serve. A signed-in session carries your own history and will flatter you.
- Several runs per prompt
- Reported as a rate rather than a yes or no. Non-determinism is why: a single check is noise.
- Verbatim capture
- What each engine actually said, kept alongside any summary score, so a disagreement can be settled by reading rather than by arguing about a dashboard.
- A competitor set
- Measured the same way on the same day. The interesting question is usually who got named instead of you.
- A failure condition
- Written down in advance: what result would make them tell you this has not worked. An agency that cannot answer this has not given you a way to fire them.
Before you trust the instrument, find out how it breaks
The failure mode that costs most · a zero that is an artefact
This is the part almost every measurement skips, and it is the one that produces confidently wrong reporting. Any tool or process you use to gather these answers has failure modes, and you have to establish them before you believe its output. Not after a result looks odd. Before.
The specific danger is a zero. A retrieval that silently returns nothing looks identical to a genuine absence: both arrive as “not cited”. A rate limit, a timeout, a blocked request, a query that quietly matched no records, a session that dropped the location. Every one of those produces a zero that means the instrument failed rather than that you were not named. Build a report on a handful of those and the trend is fiction.
The guard is cheap. Put a control in every run: a company you know for certain gets named for a given question, and a question you know returns an answer. If the control comes back zero, the run failed and the whole batch is void. Log the count of runs attempted against runs that returned anything at all, and treat any gap as an error to investigate rather than a result to report. A measurement that cannot distinguish “nothing came back” from “you were not mentioned” is not measuring visibility, it is measuring its own uptime.
Ask your agency directly: how do you know a zero in this report is real? A good answer describes a control. A blank look is the answer to a different question.
It happened to us, on this site, in a morning
Our own near miss · 12 September 2026
We set out to measure how often Google AI Overviews appear on HVAC searches.
Six queries through a widely used live-SERP service came back with no AI Overview on any of them,
including how long does a furnace last
. Six for six is a finding, and we were a few minutes
from publishing it.
Then we ran the control this page tells you to run: a question that produces an AI Overview for almost anyone. It came back empty too. That single check turned the finding into a fault, because a control returning zero means the instrument is producing the zeros.
A second source, queried for the same term in the same market on the same day, returned the AI Overview in full, with its text and all eleven cited sources. The feature was there the whole time. Our first instrument simply does not report it.
Had we skipped the control we would have published something close to the opposite of the truth, sourced and dated and completely wrong. It cost one extra query to catch. What we published instead is here, and it carries no percentage, because we never finished measuring one.
The cause, once we chased it down
The tooling was reading the wrong field. It checked for Google's classic direct-answer box, which is an older extractive feature, and treated its absence as meaning no AI Overview was present. Those are two different things on the same results page: one lifts words verbatim from a ranking page, the other writes new text from several sources. The distinction is a glossary entry on this site, and the bug was that distinction being missed inside a measurement tool.
That is the version of this worth carrying, because it generalises. The instrument was not broken in any dramatic way. It was checking a real field, correctly, and that field answered a different question from the one being asked. A tool measuring the wrong thing accurately produces confident numbers and no error message, which is why the control matters more than the tool's reputation.
What to do with the numbers
Report the rate, the denominator and the period, split by engine. Put the competitor set beside it. Then stop, because the next sentence is where most reporting goes wrong.
You may state what moved. You may not state what caused it. A model update in the same month produces a movement that looks identical to one your work produced, and there is no control group available to separate them, because you cannot run your market twice. Reporting that gives you the direction and declines to claim the cause is the trustworthy kind. The full argument for that is here.
A month of it, concretely
Fifty questions, four engines, three runs each: 600 runs, which is an afternoon with a spreadsheet. Record the answer text, the date, the engine, and whether your name appeared and whether it was linked. Two controls in every batch. Count your appearances, divide, and note every company name that came back so you have a share-of-voice figure alongside.
That produces four numbers you can defend: a citation rate, a brand mention rate, a share of voice, and a count of failed runs. The fourth is the one nobody reports and the one that tells you whether the other three are real.