Skip to content
Analytics

Measuring AI visibility: our protocol for ChatGPT and Perplexity

Short answer

Measure AI visibility with a fixed set of 15 to 30 prompts per language and answer engine. For each response, record mentions, position, tone, facts and cited sources. Those observations produce four comparable metrics: mention rate, average position, share of voice against a fixed competitor set, and citation share for your domain.

TL;DR

  • The prompt set is the instrument. Build it once, then freeze it, or you are comparing months measured with different rulers.
  • Work out the sampling error before you interpret a change. At 24 prompts and a mention rate near 40 percent, the 95 percent interval is roughly plus or minus 20 points.
  • Search Console covers Google only, and there every link inside an AI Overview inherits the same position. For ChatGPT and Perplexity there is no report at all.

Introduction

Google has had an official report since June 2026: Search Console shows impressions from AI Overviews and AI Mode. For ChatGPT, Perplexity, Claude or Copilot there is nothing comparable, and even the Google report does not answer the question clients actually ask: am I being recommended, and how often?

So you measure it yourself. A workaround, and we call it one. Below is the procedure we use at Rankprofi, written out fully enough that you can run it without us. We do not publish client numbers from these runs.

What Search Console shows and what it hides

The Generative AI performance report gives impressions for AI Overviews and AI Mode, grouped by pages, countries, dates or devices. Google is rolling it out to a subset of property owners first.

What is missing decides how useful it is. There is no query dimension: you see that a page appeared in an AI answer, not which question triggered it. You see impressions, not whether your company was named or recommended. Google excludes data from Search Labs experiments.

Three counting rules from Google's methodology documentation matter here. An AI Overview occupies a single position, and every link inside it is assigned that same position, so an average position of 1 can mean you were one of twelve source links. An impression only counts once the link has been scrolled or expanded into view, so sources further down appear less often in the statistics than in the actual delivery. And in AI Mode, every follow-up question counts as a new query.

Underneath sits a more fundamental problem. In its guide to generative search, Google describes query fan-out: the model generates several parallel sub-queries from one user question. Your impression does not map cleanly onto the question somebody typed.

Three checks before the first measurement

Before you enter a single prompt, establish whether you are allowed to take part at all. A low number is surprisingly often not a content problem but a switch someone flipped.

Google. The property settings in Search Console contain the Search generative AI control, and the default is inclusion. Set to exclude, your links and content no longer appear in AI Overviews, AI Mode or the generative features in Discover, and you receive neither impressions nor traffic from them. Google has been acting on the setting since 17 June 2026. The trap is inheritance: a property takes the value of its closest parent by default, so excluding at domain level silently excludes the blog directory too.

ChatGPT. OpenAI states that sites blocking OAI-SearchBot in robots.txt will not be shown in ChatGPT search answers, though they can still appear as navigational links. Changes take about 24 hours to register.

Perplexity. PerplexityBot is the crawler behind the results; Perplexity-User is the user-triggered fetch, which the documentation says generally ignores robots.txt rules. Check the web application firewall as well as robots.txt. A Cloudflare rule can block a bot with nothing in robots.txt to show for it, and that is the most common invisible reason for missing citations.

The prompt set is the instrument

A weak prompt set produces a worthless report no matter how carefully you analyse it afterwards. We build 15 to 30 prompts per language. Fewer, and single outliers move the percentage too far. More, and the monthly cadence stops happening. If you have to prioritise, run fewer engines with the full set rather than every engine with half of it.

The mix spreads across six intents: discovery without a brand name ("best provider for X in South Tyrol"), place and occasion ("X near Bolzano"), feature ("X with Y"), comparison ("your brand or competitor"), branded ("reviews of your brand"), and planning, the longer questions where your category is only one component. Discovery prompts are the hardest to win and worth the most.

Two rules keep the set honest. Branded prompts make up no more than about 40 percent, because they are easy hits and inflate the mention rate. And every competitor has to be reachable through at least one prompt, otherwise share of voice is systematically biased in your favour. Phrase prompts the way people actually ask, not in keyword strings nobody types.

Then freeze it. Prompts are never reworded. New ones go into a separate, dated block and enter the time series from the next period onwards. A prompt you phrased differently in May than in April is not a measurement point. It is two different questions.

For South Tyrolean clients, German and Italian run separately. The same question in the two languages produces different answers and cites different sources. Two markets in one province.

What gets logged per answer

One row per prompt and engine, seven fields.

Mentioned (yes, no, no answer). "No answer" is its own category for when the engine deflects or refuses. Those rows leave the denominator, otherwise they look like a failure when nothing was measured at all.

Position: which provider in the list you are. Blank if not mentioned, and blank means blank, not zero.

Tone: positive, neutral, negative.

Accuracy: correct, minor errors, wrong, with the specific error in the notes column. This field catches things nobody else notices: outdated opening hours, a service discontinued two years ago, a mix-up with a similarly named company.

Cited sources: every domain the engine links or names, verbatim and not normalised. The most valuable column, because it is a ready-made list of the domains you ought to appear on.

Competitors named with position where it is visible. This produces the denominator for share of voice. Plus a notes column, and as metadata the date, engine, model version, language and exact prompt wording.

Execution: a fresh, logged-out session per engine, language and region matched to the prompt, one prompt per new context. On Google the prompt is a search query rather than a chat message, and there we also record whether an AI Overview appeared at all. Those are two different findings: no overview, or an overview without you.

What never happens: adding a mention, a position or a citation because it would be plausible.

From the sheet to the metric

Mention rate is mentions divided by prompts tested, excluding the no-answer rows. Average position is the mean rank across rows with a mention, and it never stands alone: an excellent average position across three mentions is not a good result. Share of voice is your mentions divided by yours plus all competitor mentions. Citation share is how often your domain appears among the cited sources, divided by all citations. Tone and accuracy map to 0 through 1 (positive 1.0, neutral 0.5, negative 0.0) and are averaged across rows with a mention. All of it overall, per engine and per language.

That rolls up into a single value between 0 and 100, weighted 45 percent mention rate, 20 percent share of voice, 15 percent position factor, and 10 percent each for tone and accuracy. The position factor maps position 1 to 1.0 and position 10 or worse to 0.

Those weights are our decision, not an industry standard. Presence dominates because not being mentioned is the worst outcome. What matters more than the size of the weights is that they stay the same month after month. The value is comparable to itself, not to another agency's.

Cadence, variance, and why one run proves nothing

We measure monthly, in the same calendar week, in the same engine order.

OpenAI states in its own API documentation that responses are non-deterministic by default and that determinism is not guaranteed even with a seed set, because OpenAI changes model configurations. That is the API. The consumer interfaces add web search, personalisation and model changes on top.

You can calculate how much noise that implies before interpreting anything. A mention rate is a proportion, and the standard error of a proportion p over n observations is the square root of p times (1 minus p), divided by n. At 24 prompts and a measured rate of 40 percent that comes to exactly 0.1, so a 95 percent range of roughly plus or minus 20 points: somewhere between 20 and 60 percent. Run the same 24 prompts across four engines and you have 96 rows, which narrows the range to about plus or minus 10 points.

That is the generous version of the arithmetic. It assumes the rows are independent, and they are not. If your entity is weak you fail many prompts at once, and the same source shapes the answer on several engines. The real uncertainty is therefore wider than the calculated one, not narrower. Anyone selling you a three-point move as a success has not done the maths.

So we only assess from the second period, and only call a trend from the third. If a known Google core update or a model change falls inside the measurement window, we mark the period as observational. A drop concentrated in one engine that just shipped an update means something different from a drop across all engines. The second case points to a real content or entity problem.

What the number does not tell you, and what follows from it

It is not traffic and not revenue. A mention is not a session.

It is not representative of every real user question. The prompt set is an editorial sample, not a measurement of total market volume.

Position in a generated list is not a ranking. It can reorder between two runs without anything changing on your website.

We test logged out. Your actual customers are logged in, with chat history and location. Their answers look different from ours.

Share of voice depends entirely on the competitor set you chose. Remove a strong competitor and your number rises without anything having changed. That is why the set belongs next to the number.

It says nothing about cause. If the value rises two months after a publication, that is a sequence in time, nothing more.

What does follow from the pattern: a high mention rate on branded prompts and a low one on discovery prompts means the engines only know you when asked about you by name. If competitors win share of voice, work on comparison content and placements on exactly the domains in your citation column. If wrong facts appear in the answers, correct your own website and align your business profile, directory listings and Wikidata.

What explicitly does not follow, Google has written down itself: llms.txt and similar files are ignored, content does not need chunking or rewriting for AI, structured data is not a requirement, and chasing inauthentic mentions is less useful than it sounds. Building a separate page per fan-out variant violates the scaled content abuse spam policy, per the same source. A poor measurement justifies none of these.

How you get into those answers in the first place is covered in our article on how companies get named in AI answers. Which discipline owns the work is settled in the difference between SEO, AEO and GEO, and if Google is your main channel, see what AI Overviews do to Google's first page. The ongoing measurement is part of our work on visibility in AI answers. If you want to know where you stand first: free analysis.

Frequently asked questions

Can I automate this instead of testing manually?

The provider APIs let you automate a lot of it, but then you are measuring something else. The consumer interfaces of ChatGPT and Perplexity do not behave identically to the API, because system prompts, web search and personalisation sit in between. Compare API runs only with API runs.

Is this the same as share of voice in media planning?

No. In its classic sense share of voice describes an advertiser's portion of a category's advertising pressure, measured in spend or contacts. Here the same term describes the portion of brand mentions inside a defined prompt set. Same formula, entirely different denominator.

Isn't the Search Console report enough?

For Google it covers impressions in AI Overviews and AI Mode. It says nothing about whether your company was named or recommended, and nothing at all about ChatGPT, Perplexity, Claude or Copilot. Both are worth having: the report as an audited data source for Google, the protocol for everything else.

Why do I get two different answers to the same question?

Because the systems are built that way. OpenAI documents that responses can vary from request to request and that determinism is not guaranteed even with a seed. Add shifting web search results and model changes on top. That is why the unit of analysis is the whole set, not the individual answer.

Sources

  1. The Generative AI performance report (Search) shows impressions for AI Overviews and AI Mode, grouped by pages, countries, dates and devices; it is rolling out to a subset of property owners first and excludes data from Search Labs experiments. Google, Search Console Help, accessed 29.07.2026.Google Search Console
  2. Announcement of the reports. Google Search Central Blog, "Introducing Search Generative AI performance reports in Search Console", June 2026.Google Search Central
  3. An AI Overview occupies a single position and every link inside it is assigned that position; an impression requires the link to be scrolled or expanded into view; in AI Mode a follow-up question counts as a new query; Search Labs data is excluded. Google, Search Console Help, "What are impressions, position, and clicks?", accessed 29.07.2026.Google Search Console
  4. Search generative AI control: inclusion is the default; exclusion means no impressions and no traffic from AI Overviews, AI Mode and generative Discover features; Google has taken the control into account since 17.06.2026; properties inherit the value of their closest parent. Google, Search Console Help, accessed 29.07.2026.Google Search Console
  5. Query fan-out; llms.txt and comparable files are ignored by Google Search; no chunking and no rewriting for AI required; structured data is not required for generative features; chasing inauthentic mentions is less useful than it seems; building a page per fan-out variant violates the scaled content abuse spam policy. Google Search Central, "Optimizing your website for generative AI features on Google Search", published 15.05.2026, page state 08.07.2026.Google Search Central
  6. "Chat Completions are non-deterministic by default"; determinism is not guaranteed even with a seed, because OpenAI changes model configurations. OpenAI, API documentation "Advanced usage", accessed 29.07.2026.OpenAI
  7. Sites that block OAI-SearchBot will not be shown in ChatGPT search answers but can still appear as navigational links; robots.txt changes take about 24 hours. OpenAI, "Overview of OpenAI Crawlers", accessed 29.07.2026.OpenAI
  8. PerplexityBot surfaces sites in Perplexity results; Perplexity-User generally ignores robots.txt rules; changes take up to 24 hours; guidance on allowlisting in Cloudflare and AWS WAF. Perplexity, "Perplexity Crawlers", accessed 29.07.2026.Perplexity