Most AI visibility tools are precision theatre. Here's how to spot the real ones.
Imagine a dashboard tells you your brand has 35% AI visibility. It looks official. There's a trend line. Maybe a badge showing you're ahead of a competitor.
Ignore all of that for a second. The number is much softer than it looks. Treating it like a solid KPI before you understand how it was built is how marketers end up spending real money based on fake precision.
Here's why.
The same brand can look completely different depending on which AI you ask
Fractl, a research firm, recently tested how brands show up across ChatGPT, Gemini and Claude. They ran 96 prompts through each model, fifteen times, across eight industries. The result: over 8,500 brand mentions, but each AI model had its own personality. Claude favoured SaaS and insurance brands. Gemini favoured travel and healthcare brands. Only ChatGPT overlapped much with the other two.
A separate study, from a company called Boring Marketing, tested real buyer questions across four AI platforms. It found that 53% of brands never got mentioned at all.
Two different studies. Same lesson: which AI model you test matters more than almost anything else in the report. So when a tool tells you your "AI visibility," what it's really measuring is your visibility on that one model, at that one moment.
And that moment doesn't last. Data from Similarweb shows ChatGPT's share of AI traffic dropped from about 76% to about 53% in just one year, while Claude's share tripled. If your visibility score is based on last year's mix of AI tools, it's already out of date.
Why this problem exists
This isn't new. Every time a new marketing channel shows up, the same thing happens. New tools appear to measure it. Everyone uses different methods. Eventually people notice the numbers don't agree.
It happened with viewability. It's still happening with attribution. It's happening now with AI visibility, just faster, because of one added twist: the thing being measured changes every time you ask.
A search engine ranking holds still long enough to check it twice. Ask an AI model the same question twice and you might get two different answers. There's no fixed page to point to.
The IAB (a major advertising industry body) counted more than 20 companies now selling AI visibility tools, all using different methods, all capable of giving a different score for the exact same brand. Even the IAB is cautious here. Its own VP of AI, Caroline Giegerich, has said publicly that AI-search measurement isn't stable enough yet to call this a true industry standard.
That's not a small footnote. That's basically the whole problem in one sentence.
Three quick questions to test any AI visibility score
Before you trust a number, ask these three things. If a vendor can't answer clearly, treat the number as a guess with good graphic design.
1. How many questions did they actually test? The IAB says anything under 50 test questions should be treated as exploratory, not reliable. A lot of what's sold as a finished score doesn't even clear that bar. Ask how many prompts were used and how they were chosen.
2. What counts as "visible"? Being mentioned briefly, being cited as a source, and being actively recommended are three very different things. Many tools lump them all into one "visibility" percentage. That hides the number that actually matters: how often you were the answer, not just a footnote in it.
3. Can the result be repeated? Run the same test again next week. If the score jumps around with no real change in your content or the market, the tool isn't measuring your brand. It's measuring random noise.
None of this means these tools are useless. The IAB itself splits results into two tiers: "directional" (a rough signal) and "decision-grade" (reliable enough to act on). Most tools on the market today are really only the first tier, even when they're sold as the second.
Even the best framework has weak spots
The IAB's own model, called the 4 Ps, is a genuine step forward: Presence, Prominence, Portrayal, and Persuasion. But two of the four are much shakier than they sound.
Presence (were you mentioned at all) is fairly simple to measure. Prominence (where you appeared, how high up) is also reasonably solid. Portrayal (how you were described, and whether it was accurate) is much harder. It often means using one AI to judge what another AI said, and there's no agreed rulebook for what counts as a "fair" description. Persuasion (did it lead to a click or action) only tells you someone clicked. It doesn't tell you what the AI said right before that click, or whether it was even positive about your brand.
So: trust Presence and Prominence as decent early signals. Treat Portrayal and Persuasion as still mostly experimental, no matter how confident the dashboard looks.
The bigger issue: visibility might be the wrong goal entirely
Here's the deeper problem. This whole category assumes the goal is showing up inside an AI's answer.
But AI tools are moving fast toward doing things for people directly: booking a flight, comparing products, completing a purchase, all without a click or a visit to a website. If that keeps happening, "did we get mentioned" stops mattering as much as "did the AI actually pick us."
Most AI visibility tools aren't built to measure that yet, because they were designed for a world where clicks still happen. That's a real risk over the next two or three years, and it's worth thinking about now, before locking into a long vendor contract.
What to actually do
Three simple moves.
Don't sign a 12-month contract. Sign a short period one as a trial. If a vendor can't reproduce last month's score when you ask, don't renew.
Ask for the breakdown, not just the headline number. "Mentioned" and "recommended" should never be blended into one score. If a vendor won't separate them, don't trust the top-line percentage.
Tie at least one visibility number to something real. Referral traffic, branded search, assisted conversions, something. If your visibility score climbs but none of those move, that's a good story for the vendor, not for you.
AI visibility measurement isn't going away, and it shouldn't. But right now, a lot of what's being sold as a solid metric is really just an early signal wearing a scoreboard's clothing. The brands that win here won't be the ones with the highest score. They'll be the ones who figured out earliest which scores were actually real.