Visibility Score
A visibility score is a composite number that tracks how often your brand shows up when people ask AI assistants like ChatGPT, Gemini, or Perplexity about your category. It rolls up mentions, citations, and outright recommendations across a fixed set of prompts into one figure you can watch move over time. It's a tracking metric, not a ranking system — there's no shared scale across tools, so the number only means something against your own history and your own competitors.
Why the score matters
Search visibility used to be legible: you held position three or you didn't. AI assistants collapse that list into a single generated answer, and you either appear in it or you don't. ChatGPT alone counts roughly 900 million weekly users as of early 2026, Google's Gemini app has passed a billion monthly users, and Google's AI Overviews reach over two billion people a month — that's a lot of answers being handed out with no results page attached. A visibility score turns "are we showing up in that?" from a hunch into a trend line you can report on: down 22 to 41 over a quarter is something you can act on, whereas a shrug isn't. It also matters because AI citation doesn't track search rankings the way you'd expect — one analysis found only 6–8% overlap between the URLs ChatGPT cites and a site's Google top-10 rankings for the same queries. You can rank well and still be invisible to the assistants, which is exactly the gap this score is built to catch.
How it works
You start with a fixed set of prompts real customers would plausibly ask — not keywords, actual questions in natural language. You run them against several AI systems on a regular schedule and log, for each answer, whether you're named, whether you're cited as a source, and whether you're actively recommended rather than just listed. Those three outcomes get different weights, because they aren't equally valuable: an active recommendation matters far more than an incidental mention buried in a list. Averaging the weighted hits across prompts and systems gives you a single number, often expressed on a 0–100 scale, though the scale itself is arbitrary and varies by tool. What makes the number trustworthy isn't the formula, it's repeatability — the same prompts, the same systems, the same cadence, run consistently, so a change in the score reflects a change in your visibility and not a change in your test.
Common mistakes
The most common failure is an unstable prompt set: run five questions this month and twelve different ones next month, and you're measuring noise, not progress. Lock the list and keep it fixed. A second mistake is testing only one assistant — your customers spread across several, so your measurement should too. A third is treating every appearance as equal, when a stray mention in a long list is worth far less than a direct recommendation. A fourth, and increasingly common one, is chasing AI-specific technical fixes that don't move the number: Google has said explicitly that no special schema, markup, or llms.txt file is required for AI Overviews or AI Mode, and has warned against writing content "for AI" as a category. That advice is borne out in practice — one analysis of roughly 137,000 sites that published an llms.txt file found about 97% saw no measurable referral traffic tied to it, and Google's own John Mueller has confirmed no Google Search system reads those files at all. Spend the effort on being genuinely citable instead of on files search engines don't read.
Relation to AI recommendations
The score's most valuable component is the recommendation share, not the mention share. A bare mention gets you noticed; an active recommendation gets you chosen. That's why a well-built score reports the two separately — it's the only way to tell whether you're visible but unconvincing, or genuinely being put forward as the answer. Moving that ratio isn't about gaming AI systems; it tracks with being talked about elsewhere. An analysis across roughly 75,000 brands found that how often a brand is mentioned across the web correlates with AI citation rate at about 0.664 — roughly three times stronger than the correlation with backlinks, at about 0.218. In practice, that means earning genuine third-party coverage and citable, fact-checkable content moves this score more reliably than link-building does on its own.
Example
Picture a mid-sized bike shop that runs fifteen typical customer questions against three AI assistants every month, things like "where do I buy a good cargo bike in Cologne?" In January the shop turns up in only two of fifteen answers — a score of 14. The owner then builds out a clear FAQ page, publishes accurate opening hours and inventory details, and gets a few real customer write-ups picked up by local sites. By April the shop appears in nine of fifteen answers, three of them as an outright recommendation, and the score climbs to 52. It's an illustrative scenario, not a published case study, but it shows the mechanism: the number moves when the underlying facts about the business become easier for an AI system to find and trust.
Common questions
What visibility score counts as "good"?
There's no universal threshold, because scales and prompt sets differ by provider and by whoever built the measurement. What matters is the trend over time on a fixed methodology, and how you compare against direct competitors answering the same prompts. A rising score on stable measurement beats any absolute number.
How often should I measure it?
Monthly works for most businesses. AI answers vary from one run to the next even for the same prompt, so a regular cadence smooths out that noise. The frequency matters less than consistency — keep the questions, the systems, and the scoring rules fixed, or two measurements stop being comparable.