Skip to content

How many checks before an AI visibility score means anything?

Ali Khallad4 min read
September 7, 2026 , 4 min read
Share

When a brand shows up in AI answers at all, it usually doesn’t show up every time. Four out of five of those cases were inconsistent. We ran the repeat checks to see how many times you need to look before an AI-visibility reading means anything, and how often a single check misleads.

Each combination was one AI platform, one brand and one question. We checked those combinations repeatedly over 60 days. A ChatGPT check stays a ChatGPT check; a Google AI Mode check on the same question is a different combination, not another sample of the same reading. Only combinations checked at least 5 times count, which keeps two-or-three-run strays out of the set. 536 combinations qualified, averaging 8.8 checks each.

In 383 of the 536, the brand never appeared in any check. That’s most of the set, and it’s a stable miss, you can call it from a short run. 30 of the 536 went the other way, the brand appeared in every check. 123 of the 536 flipped, appearing on some checks and not others.

Of the 153 combinations where the brand appeared at least once, 123 were inconsistent. That is 80.4%. Those 153 averaged 10.4 checks each. On the inconsistent ones, the brand appeared in 41.7% of checks on average, so even the typical case where the brand shows up at all is still absent more often than it’s present.

On those inconsistent combinations a single check agrees with the majority verdict 72.2% of the time. Roughly one check in four points the wrong way. If most runs put the brand in, the next check can still leave it out, and if that was the check you kept, you’d call yourself missing. It runs the other direction too. We measured that 72.2% on the inconsistent combinations. The 383 that never appeared and the 30 that always did already line up with a single check.

One check tells you whether you appeared that time. It doesn’t tell you whether you usually appear. If a number moved and you measured once on each side, you haven’t seen a change, you’ve seen two samples of something that varies. We think a single-check screenshot of an AI answer used as proof of anything is close to worthless, and those screenshots are everywhere.

How often it flipped, by platform

The share of all checked combinations that flipped is different on each platform.

PlatformShare of checked combinations that flipped
Claude39.0%
Gemini26.6%
Google AI Mode25.0%
ChatGPT18.6%
Perplexity18.4%

Claude flips most often, and that is mostly not a fact about Claude. A combination can only flip if the brand shows up sometimes, and the brand appeared at all on about half of Claude’s combinations, against a fifth to a third of every other platform’s. More chances to be inconsistent, so a higher flip rate. Claude is also the thinnest slice here, 41 combinations at 6.2 checks each, well under the 8.8-check average of the full set. A platform that rarely names the brand looks calm on this measure, because most of its combinations never enter the flip set at all. This is not a ranking of platform stability and we wouldn’t use it as one.

When extra checks earn their keep

Run count matters most when the brand is borderline. A brand that never appears and one that always appears are both readable from very few checks. 383 combinations never appeared and 30 always did, so a large share of this set is easy to call. The 123 that flipped sit in the middle, around that 41.7% appearance rate, and the middle is where repetition earns its keep.

These are the answers a user gets from the assistant, so this is how much the number moves in normal use. We measured how much it varies, not what makes it vary.

This is a small sample, brands in competitive categories, on the questions buyers actually ask about those categories. Directional. Not a law of the web.

We measured a related thing earlier: how much the sources behind an answer turn over between checks. That was about which pages get cited. This is about whether the brand shows up at all. Repetition is what turns a reading into a measurement, and the middle cases are where it matters.