A new preprint put one of the most quoted numbers in GEO strategy under a controlled test — and that number mostly fell apart. In raw transcripts of a real search agent, the first result got cited 85.1% of the time versus 42.8% for the fifth, a 42.3-point gap that anyone tracking AI citations has probably leaned on. But when researchers Sriram Selvam and Anneswa Ghosh actually swapped the order of matched sources and replayed the conversation, the controlled effect shrank to +7.9 points in their main test and exactly 0.0 points in a held-out set of 56 pairs.
The paper, posted to arXiv on September 14 as "CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search" (arXiv:2609.15164), also found something more actionable: pages rendered with headings, short paragraphs, and lists or a table earned about half a citation more per answer than the same content rendered as plain prose — without taking credit away from the competing source. The catch, and it's a significant one, is that this moved credit between already-eligible sources. It did not clearly increase whether a page got cited at all.
What the researchers actually did
Most citation studies observe what an AI engine does and correlate it with page features. CiteChoice instead runs a counterfactual experiment on real traffic — minus the live web.
The team first prompted a GPT-5.4 search agent, using Exa as its search provider, to answer 130 everyday questions with independent web searches, recording every message and tool result. From 129 captured transcripts, they extracted 113 pairs of documents that appeared in the same search call and were both verified as supporting the same pre-specified fact — the key condition being that either page could fairly be cited, so any difference in credit comes from how the model allocated recognition, not from one page being substantively better. Blinded human review confirmed 103 of the 113 pairs.
Then came the replay. Each saved conversation was re-run four ways in a 2×2 design: the target page placed above or below its competitor, and the target's text shown either as polished prose or rewritten with structure. Nothing else in the transcript changed. An integrity gate diffed every manipulated transcript against the original and would fail closed if any undeclared field changed; all 452 trials passed. Live search was disabled, so the manipulations never touched crawling, retrieval, or ranking — only what the model sees after retrieval.
Both text versions were AI rewrites generated from the same archived evidence, with fidelity checks on claims, quantities, entities, and caveats. Because the two arms aren't word-for-word identical, the authors are careful about what they claim: the estimand is "jointly generated structured versus prose rendering," not pure formatting. To probe formatting alone, they added a mechanical ablation that took the prose version's exact word sequence and re-laid it as one sentence per list row.
Finding one: structure concentrates credit rather than winning admission
The strongest result in the paper is the citation-count effect. Structured rendering raised the target's citation count by +0.50 citations per answer (95% CI [+0.20, +0.84], Holm-adjusted p=.033), and raw means went from 2.69 to 3.12 markers per answer — a 16% increase. Distinct answer sentences citing the target rose by the same amount, so this wasn't the same sentence getting cited twice.
Where did the extra credit come from? Not from the competitor, which changed by only −0.02 markers. Not from an expanded citation budget: structured answers actually contained 0.49 fewer total markers and 0.05 fewer unique cited sources on average, none reliably nonzero. The credit concentrated on the restructured page.
But the number most GEO practitioners would care about — whether the page gets cited at all — was less conclusive. The pre-specified incidence effect was +4.5 percentage points (95% CI [−1.4, +10.4], p=.168). The design could reliably detect only effects of about 8.5 points or larger, so this is unresolved rather than evidence of zero. In raw terms: 18 families favored structured, 10 favored prose, and 61 didn't change.
The mechanical ablation complicates any simple "add lists" recipe. Over all 113 pairs, the word-preserving re-layout raised incidence by +6.7 points (p=.028) — but on the 30 repeatability families the effect reversed, landing between −1.7 and −8.3 points. The authors treat that instability as a warning rather than a mechanism. Their own framing, from the discussion section: "This is an attribution-sensitivity warning, not an optimization tactic."
Finding two: the rank gap is real in the data but mostly not in the experiment
Here's where the paper punctures a piece of received wisdom. Observational transcripts show first-exposure citation incidence falling 42.3 points from rank 1 to rank 5 among the five Exa results in a call — "rank" here means order within one search call, not a Google position. That gradient conflates position with page quality, since providers rank more relevant pages higher.
The controlled replay tells a different story. Swapping a target above its competitor produced +7.9 points (95% CI [+1.1, +14.9]), which did not survive multiple-testing correction (Holm p=.350). In 56 held-out pairs where only order was switched, the estimate was exactly 0.0 points, with a confidence interval of [−5.4, +5.4]. The observational gradient is roughly five times the scaled experimental estimate, and the experiment found no held-out confirmation of it.
The paper's conclusion is narrow but pointed: position can move attribution in some frozen transcripts, especially in wider swaps, but average rank effects inferred from observational data are not reliable. If your AI-citation dashboard implies that moving from slot 5 to slot 1 multiplies citation odds, that inference is built on a correlation that has now been tested — and didn't hold.
Finding three: single-answer citation tracking has a quantified noise floor
The third result may matter most to anyone selling or buying AI visibility metrics. When the researchers regenerated 120 answers from identical frozen inputs, the binary decision to cite the target flipped in 15% of cases — one in seven. Exact cited-pair sets agreed 78.3% of the time; exact citation counts agreed only 61.7%. Decoding randomness accounted for an estimated 45% of single-run effect variance.
That lines up with independent evidence. SparkToro's January 2026 study, which ran 2,961 prompts across ChatGPT, Claude, and Google's AI Overviews and AI Mode with volunteers repeating each prompt 60 to 100 times, found the same brand list appeared less than 1% of the time on ChatGPT and Google AI. And Ahrefs' May 2026 experiment tracking 1,885 pages that added JSON-LD schema against a 4,000-page control found no statistically significant citation increase on Google AI Overviews, AI Mode, or ChatGPT — despite the widely cited correlation that pages in AI answers are roughly three times more likely to carry schema.
Three separate teams, three different methods, one converging conclusion: a single AI answer is a noisy measurement, and anything built on one run inherits that noise.
What this means for GEO work
The paper has real limits, and they're worth stating plainly. Everything is conditional on one retrieval provider (Exa, five results per call), one model (GPT-5.4) doing the acquisition and answering, and offline replay of archived transcripts — replay can't measure whether a publisher-side edit changes retrieval or live-web ranking. Both text arms were AI rewrites, so the structure effect packages wording changes with formatting. The strict citation contract (median 29 markers per answer) may not transfer to systems that cite sparsely. And a Grok 4.3 replay was directionally positive after protocol repair, but fewer than half of its responses initially met the citation format, so the cross-model evidence supports direction, not magnitude.
Within those limits, three implications stand out for practitioners:
Treat vendor correlation claims as untested until someone swaps the variable. The schema case is the cleanest example: a 3× correlation in observational data, then no effect in a controlled experiment. CiteChoice extends the same pattern to rank position. If a tool tells you feature X drives AI citations, ask whether that claim comes from observation or intervention.
Formatting changes are real but redistributive. Structuring a page can shift citation credit toward it among sources an answer engine has already surfaced — but the evidence doesn't support reformatting as a way to break into answers where you weren't cited before. The +4.5-point incidence estimate is positive but unresolved. Anyone selling "structured content optimization" as an admission lever is overselling what this study supports.
Sample size applies to AI answers, not just surveys. One answer per prompt is a sample of one from a noisy generator. Teams tracking AI visibility should rerun prompts multiple times and report variance or visible rates rather than single-answer positions — advice that matches what SparkToro found independently and what CiteChoice now quantifies from the measurement side. The paper's own recommendation for evaluators: repeat cells, report agreement between generations, and treat single-generation studies as inheriting a quantifiable noise floor.
There's also a quieter implication for tooling. Because extraction and serialization decide what headings, lists, and table rows the model actually sees, the authors argue that document-processing pipelines are part of the attribution chain and should be versioned and audited like any other part of the stack.
CiteChoice is one preprint, not peer-reviewed, with one provider and one model, and its authors are explicit that source admission, pure serialization effects, and a general rank mechanism all remain open questions. But it's a rare example of the GEO conversation moving from correlation to intervention — and the interventions, so far, support smaller claims than the dashboards do.
Sources
- CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search (arXiv:2609.15164, posted September 14, 2026) — https://arxiv.org/abs/2609.15164
- CiteChoice HTML full text — https://arxiv.org/html/2609.15164v1
- SparkToro, "NEW Research: AIs are highly inconsistent when recommending brands or products" (January 27, 2026) — https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/
- Ahrefs, "We Tracked 1,885 Pages Adding Schema. AI Citations…" (May 2026) — https://ahrefs.com/blog/schema-ai-citations/
- OpenAI, "Introducing GPT-5.4" — https://openai.com/index/introducing-gpt-5-4/
- Exa, "Introducing Exa Agent" — https://exa.ai/blog/exa-agent
Author: Isabel Grant, Researcher of 2,000+ AI Citation Patterns at Auspia. Isabel writes about AI citation patterns, answer-engine attribution, and how document structure changes which sources get credit.




