Baseline Measurement
A baseline measurement is a documented snapshot of how often, how accurately and in what position AI assistants such as ChatGPT, Claude, Gemini and Perplexity mention your brand, taken before you change anything. It fixes a set of test prompts, a set of models, and a set of answers as the record you compare every later measurement against. Without it, "we improved AI visibility" is an opinion, not a result.
Why the baseline measurement matters
AI answers are not static: the same question put to the same model can produce different citations from one run to the next, and models themselves get updated on their own schedule. Against that noise, a claim like "our mention rate went up" is meaningless unless the starting number was captured under fixed conditions first. The baseline measurement is that fixed starting point. It also guards against the natural bias to read random fluctuation as progress, or to quietly change the questions being asked between rounds. Because AI-engine citation follows a different logic than search ranking, a study by Ahrefs found only about 6 to 8 percent overlap between URLs cited by ChatGPT and Google's top-10 results for the same query, you cannot infer AI visibility from search rankings. It has to be measured directly, starting with a baseline.
How a baseline measurement works
Start by writing a fixed list of prompts that your actual audience would plausibly type into an AI assistant, phrased the way people actually ask, not as keywords. Run each prompt against several assistants, and run each one more than once per assistant, since a single answer is a sample of one and tells you little on its own. For every answer, record whether your brand appears, at what position, in what tone, whether a source link is attached, and whether the claim made about you is even accurate. Keep the model versions, the exact prompt wording and the date fixed, because any of those changing between rounds breaks the comparison. The output is a small set of numbers, typically a mention rate and a citation rate, that becomes the reference every later measurement is judged against.
Common mistakes
Testing a prompt once and treating the answer as fact is the most common error, since AI models can and do answer the same question differently across runs. Changing the prompt wording between the baseline and a later check compares two different measurements, not progress. Not recording which model version answered is another quiet failure: a jump in mentions might just reflect a model update, not anything you did. Skipping the raw answers is equally costly, because without them you cannot go back and check what actually changed later. And treating every mention as equally good is a mistake in itself, a hallucinated or negative mention should not be scored the same as an accurate, favorable one with a working citation.
Relation to AI recommendations
Everything GEO work aims at, more frequent, more accurate, more favorable mentions in AI answers, only has meaning relative to where you started, and the baseline measurement is that starting point. It is also where you find your actual gaps: a business might be recommended reliably for one query and never appear for a closely related one, which points straight at what to work on. Third-party mentions matter here more than most people assume: an Ahrefs analysis across roughly 75,000 brands found web-mention frequency correlates with AI citation rate at about 0.664, close to three times the correlation seen for backlinks at about 0.218. A baseline measurement that only tracks your own site's changes misses that lever entirely. For client communication, the baseline is also what turns a vague promise into a checkable number: not "better visibility," but "mention rate from 15 to 40 percent by a fixed date."
Example
A bicycle shop in Leipzig wants to know how present it is in AI answers before spending any effort on it. The owner writes ten realistic questions, such as "Where can I buy a cargo bike in Leipzig?", and puts each one to ChatGPT, Claude and Perplexity five times, noting every time whether the shop is mentioned, at what position, and whether a link is included. Out of 150 total answers the shop appears in 9, always without a link. That works out to a 6 percent mention rate with zero citation rate, and that pair of numbers becomes the baseline. Six months later, running the identical ten questions against the same three assistants shows in plain numbers whether anything actually changed.
Common questions
How often should I repeat a baseline measurement?
The baseline itself is a one-time snapshot, taken once before you change anything. After that you run follow-up measurements on a fixed cadence, such as monthly or quarterly, using the exact same prompts, and compare each round back to that original starting value.
Is it enough to test just one AI assistant?
No. ChatGPT, Claude, Gemini and Perplexity pull from different sources and can answer the same prompt very differently, and each now reaches an enormous audience on its own, Gemini alone passed 1 billion monthly users in August 2026. A baseline built on a single assistant will not tell you much about your actual AI visibility, so test several and score them separately.