Quick answer: Asking Google the same question in three different ways returned almost entirely different sources. Only 14 percent of citations survived all three phrasings, and not one of our 19 questions returned the same source list across all three. When we asked the identical question three times in an hour back in July, 87.5 percent of citations survived and 17 of 20 lists were identical. The wording of the question, not the hour it is asked, is what re-rolls the answer. We call it the paraphrase re-roll.
Original study, 19 questions asked three ways, 57 live Google AI Overview pulls, 31 August 2026.
The numbers, up front
- 14.0 percent of everything cited appeared for all three phrasings of the same question (37 of 265 question-source pairs).
- 0 of 19 questions returned an identical citation list across the three phrasings.
- 63 percent of everything Google cited turned up for one phrasing only (167 of 265 question-source pairs).
- By comparison, our Citation Lottery study asked the same wording three times in one hour and got 87.5 percent survival with 17 of 20 lists identical.
- Median survival per question was 15.4 percent. The best was 50 percent. One question shared zero sources across all three phrasings.
- At domain level, ignoring which page was cited, survival was still only 17.0 percent.

What we actually did
In July we ran a study that held the wording fixed and varied the time. We pulled the live AI Overview for 20 blogging questions three times inside one hour, and found the answers were mostly frozen: 17 of the 20 came back with an identical source list every time. We called that the citation lottery, because the three that did move re-rolled 60 to 70 percent of their sources for no visible reason.
This study inverts that design. We held the time fixed and varied the wording. Same 20 seed questions. For each one we wrote two more phrasings that a reasonable person would call the same question:
- “how to start a blog” also became “steps to starting a blog” and “what do i need to do to create a blog”
- “best wordpress hosting” also became “top hosting providers for wordpress” and “which host is best for a wordpress site”
- “is blogging still worth it” also became “does blogging still pay off” and “is starting a blog a waste of time now”
All three phrasings of every question were pulled in the same session on 31 August 2026, from one Chrome profile with United States and English parameters, using the same extraction method as the July study: real browser navigation, wait for the AI Overview block, expand it, harvest the external links, reduce each to host and path.
We planned 60 pulls. Google rate-limited us at the last question and served a CAPTCHA, which we did not solve and did not work around. So the study is 19 complete question triples, 57 live pulls. A 58th pull, the first phrasing of the twentieth question, did succeed, but a triple missing two thirds of itself is not a data point, so the whole question was dropped instead of half-reported.
What happened when we changed the wording
The citation lists fell apart.
Across 19 questions there were 265 question-source pairs, which is what you get by adding up each question’s own list of everything cited for it. Thirty-seven of those pairs held for all three phrasings. That is 14 percent. (Counted globally instead, ignoring which question they came from, the study touched 232 distinct sources across 136 domains in 400 citation instances.) In the July study, using the same extraction and the same sanitising, the equivalent figure was 87.5 percent.
Put another way: the July result said the AI Overview is stable. It is stable, but only against the clock. Against the question itself it is not stable at all. Two people asking Google the same thing in their own words are served answers built from largely different source material, and nothing about the timing explains it.
And this is not a case of a stable core with a churning tail, which is what we found in July among the three re-rollers. Here the core barely exists. Median survival was 15.4 percent. The best-behaved question, “best wordpress hosting”, held 50 percent of its sources, and it is also one of the two we could not re-verify. The next best held 38.5 percent, and nothing else reached 21 percent.

Where did all the sources go?
Of the 265 question-source pairs, 167 appeared for exactly one of the three phrasings. Sixty-one appeared for two. Thirty-seven appeared for all three.
Nearly two thirds of everything Google cited was reachable through one specific way of asking, and invisible through the other two.

This is the part with a direct consequence for anyone measuring their own AI visibility. If you check whether your page is cited for one phrasing of a keyword and find that it is, you have learned something about that phrasing. On this data you have learned very little about the question.
Do any phrasings behave as the same query?
A few. Three of the 57 possible wording pairs returned an identical source list, which is 5.3 percent of pairs. Two of those three came from the two questions we could not re-verify, so the conservative version is 1 of 51 guarded pairs, or 2.0 percent. “what is affiliate marketing” and “affiliate marketing explained” produced the same nine sources in the same order. So did “how to promote your blog” and “how do you market a blog”, and “best wordpress hosting” and “top hosting providers for wordpress”.
That looks like query canonicalisation: some rewordings are recognised as the same query and served the same cached answer, and most are not. We cannot see which side of that line a phrasing falls on before we ask, and neither can you. Which is exactly the problem.
What survived the re-roll?
Thirty-seven question-source pairs held all three phrasings, and the largest single holder is awkward for a site written for bloggers. YouTube accounts for 10 of the 37, which is 27 percent of all survivors against 14.5 percent of citations overall. We are holding that loosely, because our URL sanitiser collapses every video into one key and the limitations section explains why that inflates it. After YouTube, Wix held three, Reddit and Zapier two each, and the rest were single pages: ProBlogger, BrightEdge, Thrive Themes, TechRadar, CNET, Pantheon, WordPress.org, Wikipedia, Investopedia, Salesforce, BigCommerce, Yoast, seo.com, AdSense and a handful of independent blogs.
We did not audit those 37 for word count, schema or structure, so we are not going to tell you what they have in common. What we can say is what the list is not. It is not what our schema study or our freshness study would have predicted, because neither of those signals separated cited pages from uncited ones in the first place. Working out what makes a source survive a reword is the obvious next study, and we have not run it.
Why this breaks the standard GEO playbook
Almost every piece of GEO advice in circulation tells you to optimise a page: add schema, refresh the date, write in the first person, hit a word count, structure your headings. We have now measured most of those signals on our own corpora, and the pattern is consistent and awkward.
| What we measured | Cited pages | Uncited pages | Verdict |
|---|---|---|---|
| Structured data (138 pages) | 84 percent | 89 percent | No signal, runs backwards |
| Median last update (156 pages) | 3.5 months | 4.2 months | Three weeks of difference |
| First-person language (56 pages) | 1.47 per 100 words | 1.73 per 100 words | No signal, runs backwards |
| Ranks in organic top 10 (376 citations) | 68 percent of citations do not rank top 10 | Rank does not predict citation | |
Every page-level lever we have tested is flat. Meanwhile the signals that do move are not page-level at all. They are about what kind of source you are. Reddit is on page one for 30 of 30 blogging queries and holds the top slot on 19 of them. Brand and corporate sites take 48.1 percent of AI Overview citations against 36.0 percent for independent blogs, and the brand share climbs to 61 percent on transactional queries.
Add the paraphrase re-roll to that and you get a conclusion that is uncomfortable but hard to escape. You cannot optimise a page for a query, because the query is not a stable object and the page-level dials do not turn anything. The thing that is stable is what kind of source you are and how many of the plausible phrasings you are eligible for.
That is a different game with a different strategy. Ranking is a tournament, where a small number of winners take almost everything and the correct move is to concentrate. Citation, on this evidence, behaves much more like a draw from a pool of broadly equivalent candidates, where the correct move is to be in more pools more often. Our commodity content study found that 89 percent of top-ranking blogging pages were interchangeable, which is the same conclusion arriving from the other direction: when the candidates are equivalent, the selection between them cannot be doing much work.
So what should you actually do?
Cover the phrasings, not the keyword. If nearly two thirds of what Google cites is reachable through one phrasing only, then a page written to answer one phrasing is eligible for one draw. Write the page so it answers the question in the several forms real people ask it, with those forms present as headings and as sentences, and you are eligible for several.
Stop reporting single-query citation wins. A page cited for one phrasing on one day is a coin that came up heads. Before you record a citation verdict, sample the question at least three ways. This is now the advice we give inside our AI Citation Grader, alongside the July finding about sampling across time.
Be the whole answer, not a slice of it. The 37 sources that survived all three phrasings were mostly pages with enough scope that no rewording could make them irrelevant. That is a content decision, not a markup decision.
Treat GEO dashboards with suspicion. Any tool reporting your AI visibility from one tracked phrasing per keyword is measuring one draw and reporting it as a position. On this data that number has a margin of error most dashboards do not display.
Do not spend another afternoon on schema for citation reasons. We have measured it twice now. It is a fine thing to have and it is not the lever.
What this study cannot tell you
We would rather list these than have someone else find them.
Nineteen questions is a small sample, and they are all in the blogging and SEO niche on Google US. The size of the effect here is large enough that we doubt it reverses in another niche, but we have not shown that.
Our sanitiser strips query strings from URLs, which it has to, because the browser tooling refuses to return them. That means every distinct YouTube video collapses into the single key youtube.com/watch. It understates how many videos were cited and it overstates how well video survives rewording. If you remove that key entirely, survival falls from 14.0 percent to 10.9 percent, so the headline holds either way, and the video finding above should be treated as a hint, not a result. The July study used the same sanitiser, which is why the 87.5 versus 14.0 comparison is like for like.
Two of the 19 questions were captured before we added a guard that checks the harvested page really is the query we asked, and rate limiting stopped us re-running them. Those two are “how to promote your blog” and “best wordpress hosting”, and both show up above: between them they supply two of the three identical wording pairs and the top bar on the per-question chart. Dropping both moves the headline from 14.0 percent to 12.9 percent, the identical-pair rate from 5.3 percent to 2.0 percent, and the best question from 50 percent to 38.5 percent. The headline does not depend on them. Those two smaller claims partly do, which is why they are flagged where they appear.
Survival is capped by the shortest list. Pull sizes ranged from 2 to 12 sources. Because survival divides the sources common to all three phrasings by the union of all three, a question whose smallest pull returned two sources cannot score highly however stable it is. Given the list sizes we actually observed, the highest survival this study could have recorded is 39.2 percent, not 100 percent. Two things follow. Per-question rates are not strictly comparable with each other, and the gap to the July baseline is somewhat flattered, because in July the three lists were mostly identical so the ceiling there sat close to 100 percent. The direction of the finding survives this. Part of its size does not. If you want the same data in a measure with no ceiling problem: take any one phrasing and compare it against another phrasing of the same question, and 57 percent of the first one’s cited sources are missing from the second.
Three pulls per phrasing set is a floor, not a law. And we compared against a July baseline, so a small amount of the gap could be five weeks of drift and not wording at all. We think that is a small part of a very large gap, given the July study found the same-wording answers frozen within the hour, but we cannot separate the two from this design.
We did not solve the CAPTCHA. The twentieth question is simply missing.
Frequently asked questions
What is the paraphrase re-roll?
The paraphrase re-roll is what happens when the same question is put to Google’s AI Overview in different words: the cited source list is largely regenerated instead of reused. In this study only 14 percent of sources survived across three phrasings of the same question, against 87.5 percent when the wording was held constant.
Does this mean AI Overviews are random?
No. They are highly stable against time and highly unstable against wording. Asked the identical question three times in an hour, 17 of 20 AI Overviews returned an identical source list. The instability is triggered by the phrasing, which suggests per-phrasing caching and retrieval, not randomness.
Should I write a separate page for every phrasing of a keyword?
No. That is the old keyword-page reflex and it produces thin, cannibalising pages. The better move is one page with enough scope to answer the question in its several natural forms, with those forms present in the headings and the body.
Does schema markup help with AI citations?
Not on our data. Across 138 pages drawn from 18 live AI Overviews, 84 percent of cited pages carried structured data against 89 percent of uncited pages. The gap runs the wrong way and is not meaningful.
How should I measure my AI search visibility given this?
Sample every question at least three ways, and at least three times, before recording a verdict. Report a rate across phrasings, not a position for one phrasing. A single-query citation check is a single draw.
How was this data collected?
Nineteen blogging questions, each asked in three semantically equivalent phrasings, pulled from live Google AI Overviews in one session on 31 August 2026 from a single Chrome profile with United States and English parameters. Fifty-seven pulls, 232 distinct sources across 136 domains, giving 265 question-source pairs. The extraction method matched our July Citation Lottery study exactly so the two are comparable.
Cite this data
| Citation survival across 3 phrasings | 14.0 percent (37 of 265 question-source pairs) |
| Citation survival, same wording, 3 pulls | 87.5 percent (189 of 216) |
| Questions with an identical source list | 0 of 19 |
| Sources cited for one phrasing only | 63 percent (167 of 265 question-source pairs) |
| Domain-level survival | 17.0 percent (39 of 229) |
| Median survival per question | 15.4 percent |
| Sample | 19 questions, 3 phrasings each, 57 live AI Overview pulls |
| Collected | 31 August 2026, Google US |
Plain citation: Blogging Titan (2026). The Paraphrase Re-roll: Google’s AI Overview Cites Different Sources When You Reword the Question. bloggingtitan.com
Data licensed CC BY 4.0. Attribution to Blogging Titan with a link to this page.
- The Citation Lottery, the same-wording baseline this study is measured against.
- Schema Markup Will Not Win You AI Citations, 138 pages, no signal.
- The Rank-Citation Disconnect, 68 percent of AI Overview citations do not rank in the top 10.
- Google AI Overview versus Perplexity, the two engines agree on a source 16 percent of the time.
- The AI Citation Grader, our free tool for scoring a page on citability.