Three brass tuning forks of different lengths on a slate slab, one blurred from vibration.

There is no single AI visibility number

Only about 11% of domains cited by ChatGPT are also cited by Perplexity, so one blended score averages systems that fundamentally disagree.

Any tool reporting one AI visibility score is averaging systems that do not agree with each other. Roughly 11% of the domains cited by ChatGPT are also cited by Perplexity for the same questions. Median unique domains cited per answer runs about 6.4 for Perplexity, 3.1 for ChatGPT and 2.4 for Gemini.

Different retrieval, different indexes, different citation appetite. A blended number hides which engine you are winning and which you are absent from, and those are the only two facts that would change what you do next.

Why the engines disagree so completely

Each system builds its answer from a different pool.

Perplexity cites widely, pulling six or more distinct domains per answer on average, which makes it the easiest engine to appear in and the least concentrated.

ChatGPT cites fewer sources per answer and leans differently depending on whether search is invoked at all. Winning here is closer to a shortlist competition.

Gemini and Google's AI features run on the Google index, so they inherit Google's view of authority, and they cite the fewest domains per answer of the three. That makes Google's surfaces the hardest to enter and the most valuable when you do.

None of them are ranked lists. They are generated answers, which means the same prompt run twice can return different sources. That variance is not a measurement error to be cleaned up. It is a property of the system, and any score that does not account for it is reporting noise as change.

What a real measurement design looks like

Build it yourself, or make sure whoever supplies it can show you these five things.

A fixed prompt set. Fifty to one hundred real buying questions, written the way a buyer would ask them, not keywords. Fixed means fixed. Adding prompts mid-quarter breaks the series and every comparison after it.

Per-engine tracking. Separate results for ChatGPT, Perplexity, Google AI Overviews, AI Mode, and Gemini if it matters to your market. Never blended. If a tool cannot break it out, it is not measuring, it is scoring.

Repeat runs per prompt. Each prompt run three to five times per cycle, because a single run tells you almost nothing. Report the rate at which you appear, not a binary yes or no.

Position within citations. Being the first source cited is not the same as being the ninth. Record it.

Share of the citation pool. For each prompt, who else was cited. That is your competitive picture, and it is the part most reporting drops. Knowing that a supplier's technical page is cited on eight of your prompts is more actionable than any score.

Run it monthly. Same prompts, same method, same person. The absolute numbers matter far less than the direction.

Connecting it to money

Citation is a leading indicator, not revenue. Two things tie it back.

Referral traffic by source. AI-referred sessions are small. For most sites they sit around or below 1% of total. They also convert unusually well. Semrush measured AI-referred visitors converting at roughly 4.4 times the rate of ordinary organic. Seer's client data has shown ChatGPT-referred traffic converting near 16% against about 1.8% for Google organic. ChatGPT referral volume grew 206% year on year from January 2025 to January 2026.

Small pool, high intent, growing fast. Track it as its own channel in GA4 rather than letting it collapse into direct or other referral. Check that your referrer parsing is actually capturing the AI hosts, because many default configurations do not.

Brand search volume. If you are being cited during research, some of those people search your name later rather than clicking through. A rise in branded search alongside rising citation rate is the pattern to watch. It is imperfect. It is also the only bridge between an answer you cannot track and a buyer who arrives without a source.

What to ignore

Any composite AI visibility index. It combines engines that agree 11% of the time.

Month-on-month movement under about five percentage points. That is inside normal variance for a non-deterministic system. Reporting it as a trend will send your team chasing noise.

Tools that will not tell you their prompt set. If you cannot see the questions, you cannot know whether the score reflects your commercial reality or a list of easy queries chosen to look good.

AI Mode share panic. AI Mode is around 0.34% of searches. It is strategically interesting and commercially small. The traffic pressure is coming from AI Overviews on ordinary results pages, which is where zero-click reached 68.01% of US searches in early 2026.

What this costs

An honest measurement programme is a few hours a month of structured work, or a modest tool spend plus oversight. That is the real number, and it is small.

The larger cost is the discipline of not changing the prompt set when the numbers look bad, and not celebrating a five point move that means nothing. Marketing reporting culture rewards a single number that goes up. This discipline gives you five numbers that sometimes go sideways.

The one thing to build first

Before any tooling, write the prompt set. Fifty questions your buyers genuinely ask, sourced from your sales team and your support inbox, not from a keyword tool. That artefact is useful whether or not you ever measure with it, because it doubles as the content brief for the next twelve months.

Everything in AI search measurement depends on asking the right questions first, which is also true of the search itself.

Written by David Eid. Published .