How to actually measure AI search visibility
You can measure whether you appeared. You mostly can't measure whether it worked. Four sources cover the first half: Search Console's generative AI reports, Bing Webmaster Tools' AI Performance report, GA4's AI Assistant channel and your own server logs. None of them connects a citation to a deal. Set all four up, then learn which sentence each one earns you.
Category
Measurement
Reading time
6 min

You can measure whether you appeared. You mostly can't measure whether it worked. Four sources cover the first half: Search Console's generative AI reports, Bing Webmaster Tools' AI Performance report, GA4's AI Assistant channel and your own server logs. None of them connects a citation to a deal. Set all four up, then learn which sentence each one earns you.
Desses sells an AEO audit and measurement setup is one of the four things inside it, so recommending it is billable work for me. It gets you presence data. Attribution stays out of reach, and every vendor selling it to you is missing the same piece I am.
Two definitions, because they get swapped constantly. A citation is a link an AI system attaches to an answer it generated. An AI referral is a human clicking that reference and landing on your site. The first happens thousands of times more often, and most confident claims in this field come from someone treating them as one event.
What each of the four sources tells you
Source | What it shows | What it can't show | Cost |
|---|---|---|---|
Search Console, generative AI reports | Impressions by page, country, device and date, across AI Overviews, AI Mode and generative AI in Discover | Google's own list names no query dimension, no position and no click metric. Standard 1,000-row cap | Free |
Bing Webmaster Tools, AI Performance report | Total citations, average cited pages per day, page-level citation activity, trends, and grounding queries | No ranking, no prominence, no clicks | Free |
GA4, AI Assistant channel | Sessions and on-site behaviour from recognised AI chatbot referrers | The recognised-referrer list is unpublished. Referrer-less traffic lands in Direct. AI Overview clicks arrive as Google organic | Free |
Server logs | Which bots fetched which URLs, when, and what status code came back | Anything that happened after the fetch. No link to any answer | Your own hosting |
Three of the four cost nothing. The fourth costs a log export and an afternoon. No fifth source closes the attribution gap.
Search Console shows you appeared, then stops
Google announced the generative AI performance reports on 3 June 2026 and had them out globally by 31 August 2026. The help page lists impressions by page, country, device and date, across AI Overviews, AI Mode and generative AI in Discover.
Now read that list for what's absent. No query dimension. No position. No click metric. Google never states they're excluded, so what you have is an argument from omission, and it's the only argument on offer.
Impressions by page still earns its place. It tells you which URLs Google's generative surfaces pull from, and that's the input to every content decision after it. It's a presence log, and nothing in it measures performance.
Bing shows you the question, and almost nobody looks
Bing Webmaster Tools opened its AI Performance report in beta on 10 February 2026: citations, cited pages per day, page-level activity, and grounding queries, the sample phrases that retrieved your content before it got cited.
That last one is the only place any provider shows what was asked. Microsoft states the limit plainly: the report counts citations and nothing else. ## GA4 will undercount you and there's no fixing it
GA4 added an "AI Assistant" default channel group on 14 May 2026, auto-assigning the medium ai-assistant to recognised chatbot referrers. Google hasn't published that list, so you can't audit what it caught or missed.
The bigger hole is structural. Referrer-less traffic lands in Direct, alongside in-app browsers and copied links, and a large share of AI chat happens inside an app. Every published AI referral figure is an undercount of unknown size. AI Overview clicks arrive as ordinary Google organic and can't be separated out at all, so anyone handing you a clean AI Overview click number is handing you a model.
Run the native channel and a manual regex side by side, as a referrer exploration:
The two rarely match. That gap is your list-coverage problem in one number, and it's the most useful thing GA4 gives you.
Server logs are the only place you learn whether the bots arrive
Everything above measures what happened after retrieval. Server logs tell you whether retrieval happened at all, and on most sites I look at, that's where the real problem sits.
Pull 30 days of raw access logs and filter for the AI crawler user agents. Then the step almost everyone skips: verify by published IP range rather than the user-agent string, because spoofing one takes a single curl flag. Anthropic publishes verified ranges at claude.com/crawling/bots.json. Then separate the training bots from the retrieval bots, because that distinction decides whether any of this matters.
User agent | Operator | What it does | What its absence means |
|---|---|---|---|
OpenAI | Training corpora | Nothing for AI search | |
OpenAI | The index ChatGPT cites from | ChatGPT search has no current copy of your page | |
Anthropic | Training | Nothing for AI search | |
Anthropic | Search result quality | You're outside Claude's search | |
Perplexity | Search index, explicitly not training | You're outside Perplexity | |
Search index, feeds AI Overviews and AI Mode | Everything is broken | ||
| Microsoft | Bing index, feeds Copilot | Everything on Bing is broken |
Watch the status codes as closely as the hits. A run of 404s to retrieval bots is a crawlability job, not a measurement one, and the crawler article covers what to do about it.
The prompt set, and how many times to run it
Twenty prompts. Real ones, in the words your buyers type. Run them across ChatGPT, Claude, Perplexity and Google AI Mode, and repeat each one seven or eight times before you let any result count as signal.
The repetition carries the whole method, and one paper pays for it. Schulte et al. tracked four engines for 45 to 46 days in early 2026 and measured how much of one day's cited sources survived into the next. A Jaccard value of 0.35 means about 35% of the cited sources overlap between two consecutive days. Same engine, same question, one day apart.
An earlier draft of this page pointed that finding at the wrong thing, at overlap between engines. Fact-checking caught it before publication, and what Schulte measured makes the harder case for repeating a prompt. Cross-engine agreement has its own number: Grossman et al. put it at under 0.2 average Jaccard similarity across 7,439 queries.
Twenty prompts across four engines at eight repetitions is 640 responses. Record whether your brand was named, whether a link to your domain appeared, which competitors were named, and which domains got cited. An afternoon in a spreadsheet, once. Re-run at weeks 2, 6 and 12.
Profound, Peec AI, Otterly, Semrush's AI toolkit and Ahrefs Brand Radar all automate a version of this, and they share one flaw: they sample prompts rather than observing real user queries. Your spreadsheet has the identical flaw, costs nothing, and you know what went into it.
How to state a number so it survives a skeptic
Here's one of ours, written the way this industry usually writes them. ChargeBlock's new site produced 40% more AI search visibility.
As written, that sentence is worth nothing to you. It doesn't say what was counted: citations, impressions or appearances in a prompt set. No starting point, no window, no named engines, no source.
The version that survives a skeptic has a fixed shape. Across a prompt set of N prompts run 8 times on named engines, the brand appeared in X% of responses in month B against Y% in month A, measured in one named source, with the launch as the only deliberate change in the window.
I can't complete that sentence. The baseline was never recorded, which is why the claim doesn't appear anywhere else on this site. That's the honest state of most 40% claims in this field, ours included, and it's why measurement goes in on day one of a build.
When to skip most of this
Under roughly 50 organic visits a day, your sample is too thin for any of the four sources to say anything. Fix demand first.
If your logs come back showing OAI-SearchBot and Claude-SearchBot never arriving at all, stop measuring and fix crawlability. There's nothing to track yet.
And don't move budget out of working search to fund this. The gap between the two channels is not close. Keep the one that pays for everything and instrument the new one for free.
What we'd set up, concretely
Day one: Search Console's generative AI report exported, Bing Webmaster Tools verified (most clients we onboard never have), the GA4 AI Assistant channel confirmed alongside a manual regex exploration, 30 days of logs pulled and IP-verified. Then the prompt sheet, with calendar entries for weeks 2, 6 and 12 booked the same day, because the re-run is the part that gets forgotten.
On Framer, one open question. Every page is pre-rendered to HTML on Framer's servers, so ?md shows you what a non-rendering crawler receives. Raw server logs are the gap: Framer exposes no log export I've found, so the log half of this setup needs a proxy in front of the site, and that's a decision to take before launch.
The setup is a day of work and it belongs at the start of a build. The ChargeBlock number above is what happens when it goes in late. It's part of how we scope our website design work.
Send me your domain and I'll tell you which of the four sources is already switched on, which one you've never opened, and what your logs would need to show before any of it means anything. If you don't need this yet, I'll say that too.













