Ask an AI assistant the same buyer question two days running and it may cite a different set of sources each time. Most people who watch AI search already sense this. What has been missing is a number: how much do the AI citations actually change between one check and the next, and does it depend on which assistant you ask or on what you ask it. We logged more than 20,000 answers to a fixed set of real buyer questions across five assistants and counted.
How we measured it
We took a set of buyer questions we track, the kind someone types when they are close to choosing, and asked each one on repeat across ChatGPT, Claude, Gemini, Perplexity, and Google AI Mode. Every answer records the sources the assistant cited. For each question on each engine, we put the checks in time order and compared each check’s set of cited domains against the one right before it. The number we report is turnover: the share of cited sources that changed from one check to the next, averaged over hundreds of check-to-check comparisons per engine.
Two honest notes on scope. We measured the sources an answer cites, not whether the recommendation itself changed. Those are different questions, and the gap between them turns out to matter, so we come back to it below. And the questions span several buyer categories, on prompts where you would expect the brands to show up, so this is a benchmark sample rather than a random crawl of the web. Read the levels as directional.
The sources move more on some engines than others
Turnover was not uniform. On Google AI Mode, about two thirds of the cited sources changed from one check to the next. On Perplexity, about a quarter did. The other three landed in between.
| Assistant | Cited sources that changed between checks |
|---|---|
| Google AI Mode | 68% |
| Gemini | 48% |
| Claude | 45% |
| ChatGPT | 35% |
| Perplexity | 25% |
Perplexity held its sources steadier than anything else we tested. Google’s AI surfaces were the least stable of the five. If you check your presence on one engine on a single day and read it as the state of the world, Google AI Mode is where that reading misleads you most.
Broad questions churn, narrow ones hold
The engine is only half of it. How wide the question is mattered as much.
A narrow question, one that names a single specific capability, tended to be stable. One such question held the same top cited source in roughly seven of every ten checks on ChatGPT, and drew on only a couple dozen distinct domains across all of its runs.
A broad question, the “best tools for X” kind, churned hard. One broad question kept its top cited source in about one check out of five. Across repeated checks on a single engine, it pulled citations from more than a hundred different domains. Same question, same assistant, over a hundred sources deep in the pool it was drawing from.
I expected the broad questions to be messier. I did not expect the spread to be that wide. A hundred-plus sources for one question is a different kind of unstable than a top source that swaps every fifth check, and the two need different responses.
What churned was the sources, not always the pick
This is the gap we flagged earlier. Source churn is not the same as recommendation churn. On one narrow question, I read the actual answers across its checks, and the assistant kept recommending the same product at the top while the sources cited underneath that recommendation shifted from check to check. The citation list was busy. The answer on top of it was steady.
That is a spot check on one question, not a measured rate, so I will not put a number on it. It still changes how you read the churn. A moving source list does not prove the recommendation moved, and a steady recommendation does not prove your sources are safe. They are two separate things, and watching only one of them will surprise you eventually.
What this does not show
We did not isolate why the sources move. Location, a model update mid-window, time of day, and the assistant simply sampling differently on each run could all feed the churn, and this measurement does not separate them. The sample also skews toward a few buyer categories, so the exact percentages are directional. The ranking of engines is the more durable result. The distance between Perplexity at a quarter and Google AI Mode at two thirds is wide enough that sample noise is unlikely to flip it.
What to do with this
A single reading of a single AI answer tells you what one engine said on one run, not what it says across many. We have made that argument in general before, that one AI answer is not AI search visibility. These numbers are the size of the effect. So pick the questions that matter to you, check them on a schedule, and read the trend across checks instead of any single day. Weight the volatile engines accordingly: a one-day reading on Google AI Mode carries far less signal than the same reading on Perplexity.
The sources that matter are the ones that keep coming back. A domain cited in most checks of a question is part of that question’s stable core, and that is where earning a mention pays off. A domain that shows up once and vanishes is the pool coughing something up for a day. The turnover figure is close to the ratio of that noise to the core, and it runs higher than a single-snapshot look will ever show you. For the companion finding, how little the five engines even agree with each other on the same question, see our engine-overlap study.



