The short answer
In July 2026, a researcher named Olivier Martinez published a critical survey on arXiv called Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026). It went through 45 studies published since the GEO field began, plus related work on retrieval and evaluation, and stress-tested their claims.
Strip out the academic language and the bottom line is short:
- The famous "GEO boost" numbers are real but conditional. The widely shared 30–40% visibility gains from the original 2023 GEO study only held in that study's own experiment, and only for content the engine had already retrieved. They never proved that GEO makes you discoverable in the first place, and they never proved durable traffic.
- Two levers survive the evidence review: topical relevance and position within the AI answer's context.
- Generic "GEO checklist" tricks transfer poorly between sites, markets, and engines. Some citation-hunting rewrites actively hurt retrieval.
- No technique yet shows a stable, repeatable, cross-platform effect on organic discoverability or downstream clicks. What is proven: once your content is inside the answer's context, you can meaningfully change whether and how you are cited and used.
For practitioners, this is not bad news. It is a clearer map. Stop treating GEO as one ranking to beat. Fix the specific stage where you actually lose, and measure four separate outcomes instead of one fuzzy "visibility" number.
What this paper is (and why a critical survey matters)
Most of the GEO advice you read online traces back to a handful of studies. The 2023 GEO: Generative Engine Optimization paper is the big one, but there are also many smaller experiments, blog "case studies," and audits since. Most people read the summary headline, not the method.
A critical survey is different. It collects a whole body of research and asks: which of these results would survive if someone else ran the experiment again? That is exactly what this paper does for GEO.
What it reviewed:
Item | Detail |
|---|---|
Studies covered | 45 GEO studies, November 2023 – July 2026 |
Plus | Related work on RAG, retrieval, and evaluation |
Source | arXiv:2607.14035, published July 15, 2026 |
Author | Olivier Martinez |
Format | 18 pages, 8 tables, 1 figure |
The survey also ships its own search protocol and a literature matrix as downloadable files, so you can see exactly which studies were included and how they were screened. That transparency is rare, and it is what makes the survey worth reading even if you never touch the 45 papers yourself.
GEO in plain English
Before we go further, one quick definition so we are all on the same page.
Generative engines are the AI systems that answer questions in full sentences instead of showing ten blue links. ChatGPT, Gemini, Perplexity, Copilot, and Google's AI Overviews are the ones most of us meet daily.
Generative Engine Optimization (GEO) is the practice of making your content more likely to show up, get cited, and actually influence the answer in those engines.
The paper puts it more precisely: GEO aims to increase your content's presence, citation likelihood, or influence in generative engine answers. Notice that "presence" and "influence" are different things. That distinction is the heart of this survey.
Why GEO is a nine-stage pipeline, not a single ranking
Here is the paper's core argument, and the most useful idea for beginners to walk away with:
GEO is not a single ranking task. It is a stochastic, partially observable pipeline.
Two words there matter.
- Stochastic means random in a way you can't fully predict. Ask the same question twice and the engine may retrieve different sources. This is why a single screenshot of "your GEO result" is close to meaningless.
- Partially observable means you can't see most of the pipeline. You see the final answer, but not which candidates were considered, how they were reranked, or why your page was dropped at stage three.
The survey breaks the pipeline into nine stages. Here they are in plain English, with the typical way a normal site fails at each one:
Stage | Plain English | Where a normal site loses |
|---|---|---|
1. Search activation | A user asks something that makes the engine decide to consult the web at all | Your niche is answered from the model's memory; no web search is triggered |
2. Crawling & indexing | The engine's crawler finds and stores your page | Page blocked, slow, or too obscure to be discovered |
3. Retrieval | The engine pulls candidate sources for the answer | You are topically relevant but never make the candidate pool |
4. Reranking & context allocation | The engine decides which candidates actually enter the answer's context | You enter context, but get squeezed into a minor role |
5. Citation | Your source is named or linked in the answer | The answer uses your facts without crediting you |
6. Prominence | How early and visibly you appear in the answer | You are cited in a spot nobody reads |
7. Factual absorption | Whether the answer actually adopts your fact | The answer paraphrases you wrong or drops your point |
8. Fidelity | Whether the answer represents you accurately | The answer misstates your product, price, or position |
9. User behavior | Whether the user clicks through or takes action | You are cited but get zero clicks |
Most "GEO tools" and most GEO advice only touch stages 4–6. If your problem is at stage 2 (the crawler never found you) or stage 3 (you never get retrieved), polishing citation formatting will do nothing. The pipeline tells you where to look first.

A useful mental image: being included in an AI answer is like getting invited to a dinner party. There are many steps before you even get a seat. The host has to know you exist, decide you're worth inviting, and seat you at the table. Once you're seated, the conversation (citation and absorption) matters a lot. But the entire "GEO tactics" conversation is about what happens after you're seated. The survey's uncomfortable finding is that nobody has proven how to reliably get a seat in the first place.
The famous "GEO boost" numbers, decoded
The original 2023 GEO study is the reason the field exists. It reported that applying GEO treatments (adding citations, statistics, fluent structure) lifted average visibility in generative answers by a large amount. Marketing material everywhere still repeats that number as proof that "GEO works."
The survey does not call that study wrong. It calls it narrowly scoped:
The foundational paper's gains are valid within its experimental setting, but conditional on a source already being present in a fixed context. They establish neither organic discoverability nor durable traffic effects.
Let's translate. In the 2023 experiment, the researchers fed a fixed set of sources into the engine's context, and then tested whether small changes to those sources changed citation behavior. It is a test of citation optimization for content that was already retrieved. It is not a test of getting discovered.
So the correct reading is:
- "I am already in the answer pool → optimizing my content can measurably improve how I'm cited and used." Supported by evidence.
- "Doing GEO will get me into the answer pool in the first place." Not yet proven by any reviewed study.
The survey also reports what commercial GEO audits consistently find:
- Low source overlap. Different engines pick almost entirely different sources for the same question.
- Substantial run-to-run variability. The same engine, same question, different day, different sources.
- Persistent fidelity gaps. Answers still quote or paraphrase brands inaccurately even when the source is cited.
This is why an audit screenshot from a single engine on a single day tells you very little. It is a data point, not a result.
What the evidence says works (and what doesn't)
The survey's evidence review is its most practical contribution. Here is the plain-English version of the "levers" it tested:
Lever | What the evidence says | What to do about it |
|---|---|---|
Topical relevance | Most reproducible lever across studies | Make sure your page is genuinely, deeply about the question being asked, not adjacent to it |
Context position | Most reproducible lever across studies | Design pages so the key answer sits early and clearly in the content the engine reads |
Generic heuristics ("add stats, add FAQs, add quotes") | Transfer poorly between sites and markets | Treat checklist advice as a hypothesis, not a guarantee; test it on your own prompts |
Citation-oriented rewrites | Can impair retrieval | Don't restructure content purely to win citations. You can lose the discoverability you already had |
Competition | Erodes individual gains | What works when you're alone in a niche may do nothing in a crowded one; gains are relative |
Long-term, cross-platform effect | Not demonstrated for any single technique | Set expectations accordingly. This is a compounding content and authority play, not a switch |
The pattern worth noticing: the levers that hold up are boring. They are about what you talk about and where the answer sits. The levers that fail are the exciting ones: the trick-y, hack-y "make the AI cite you" moves.
The four outcomes you should track separately
The survey introduces a visibility vector with four components. This is the single most useful framework in the paper for someone who runs SEO or content:
Outcome | The question it answers | Example of a failure here |
|---|---|---|
Discoverability | Can the engine find and retrieve you at all? | Your brand never appears in any candidate source |
Citation | When the engine uses web sources, are you named? | The answer uses your data but links to a competitor |
Absorption | Is your actual fact or claim used in the answer? | You're cited, but the answer says the opposite of your page |
Economic outcomes | Does exposure turn into clicks, visits, or leads? | You're cited prominently, and no one clicks |

Most dashboards collapse these four into one "AI visibility score." That is like reporting a company's revenue and its profit as the same number. The four outcomes fail for completely different reasons and need completely different fixes:
- Low discoverability → fix crawling, indexing, topical architecture, entity clarity.
- Low citation → fix evidence quality, source markup, whether your page reads like an authoritative source.
- Low absorption → fix whether your claims are stated plainly, early, and accurately.
- Low economic outcome → fix the page itself. The answer may be fine and the landing experience broken.
Once you separate them, the diagnosis gets obvious. A brand that is "cited but wrong" has an absorption and fidelity problem, not a discoverability problem. Spending that month on link building would be wasted effort.
What this means for SEO and GEO practitioners
Here is the honest, actionable reading, and it aligns with what we see running AI-search visibility checks across brands:
Keep doing this. Build genuine topical relevance. Put your core answer early and clearly. Make your facts plain, sourced, and correct, so that when an engine does use you, it absorbs you accurately and shows fidelity. Establish consistent entity facts (what you are, what you sell, who you serve) so engines can attribute information to you in the first place. These are the levers the survey's evidence actually supports, and they are the same boring fundamentals that already drive good SEO.
Stop doing this. Expecting a quick, durable "GEO ranking." Copying a generic optimization checklist and assuming it transfers. Rewriting pages purely to chase citations (the survey flags that this can degrade retrieval). Trusting a single engine snapshot or a single-session audit as a verdict.
Treat GEO as pipeline work, not page work. Before you polish stage 5 (citation), check whether you're winning stage 3 (retrieval). Most brands we audit fail at discoverability, not citation. The tools to check that are simple: ask your buyer's questions across engines and see whether your content is even in the conversation. That check matters more than any rewriting technique.
Set honest expectations internally. The evidence supports improving citation and use for content already in play. It does not yet support promises of predictable organic discoverability or traffic growth. Anyone selling you "guaranteed AI rankings" is selling the part of the pipeline the research cannot yet back.
A beginner's measurement routine for AI visibility
You do not need a paid tool to start measuring like the survey recommends. Its protocol is built on five ingredients: repeated measurements, paraphrases, controls, human validation, and multi-actor comparison. Here is a 5-step routine anyone can run in about an hour a week.
Step 1: Build a prompt set. Write 10–15 questions your actual buyers ask, in their words. Not search keywords: questions a person would type into an AI engine.
Step 2: Record a baseline for each. For every prompt, note: Does my content appear at all? Is my source cited? Is my actual fact used? (The four-outcome vector from above.) You can track this in a simple spreadsheet.
Step 3: Repeat and paraphrase over several days. Ask the same questions again, and ask slightly reworded versions, on different days. This is the "stochastic" check: one day's answer is a sample, not a result. A free way to start is Auspia's AI Search Visibility Checker, which records per-prompt visibility across engines so you can compare runs instead of trusting a single screenshot.
Step 4: Compare across engines. Record the same prompts in two or three engines. With source overlap as low as the audits report, an engine that never cites you may simply be a different engine, so you need the cross-engine picture before you change anything.
Step 5: Score by hand and decide. After a few weeks, look for the pattern across your four outcomes, not any single answer. If discoverability is the gap, invest there. If absorption is the gap, rewrite for clarity. Then change one thing at a time and re-run the same prompts. That loop, measuring, changing one lever, and measuring again, is the practical version of the survey's entire methodology.
FAQ
Does this mean GEO doesn't work? No. It means the proven part of GEO is narrower than the marketing. Once your content is retrieved, optimizing it measurably changes whether and how you're cited and used. What's unproven is reliably getting retrieved in the first place. Both are worth working on, just with different expectations.
Are the "30–40% visibility boost" claims false? Not false, just misapplied. Those numbers came from an experiment where sources were already inside the engine's context. They measure citation behavior for retrieved content, not organic discoverability. Treat them as "citation gains, conditional on retrieval," not as a guarantee of ranking or traffic.
What's the difference between being cited and being absorbed? Cited means the engine names you as a source. Absorbed means the answer actually used your fact. You can be cited while the answer contradicts you. That's a fidelity problem. Measuring both separately is how you notice the difference.
Should I stop optimizing content for citations? Keep earning citations, but don't restructure content purely to chase them. The survey notes citation-oriented rewrites can impair retrieval, which is the exact thing you need to be cited at all. Optimize for being a clear, correct, topically relevant source, and treat citation as the by-product.
How do I know if I'm discoverable? Ask your buyers' questions across two or three engines over several days. If your content never shows up as a source, you have a discoverability problem: fix crawling, indexing, topical relevance, and entity clarity before you worry about citation formatting.
Is a GEO score from an audit tool trustworthy? As a diagnostic, yes; as a verdict, no. Because results vary run-to-run and engine-to-engine, a single score is a data point. The useful signal is the trend across repeated, paraphrased, cross-engine measurements, exactly the protocol the survey recommends.
Author: Iris Campbell, Editorial Evidence Analyst with 2,500+ Sources Reviewed at Auspia. Iris writes about research synthesis, evidence quality, and separating what's proven from what's hyped in AI search and GEO.












