A claim has been moving through search teams this month with a very confident number attached to it. Put Brand A first in a comparison query, and Google's AI Overview recommends Brand A roughly 80% of the time. Reverse the order, and the recommendation reverses with it. Size does not save the bigger brand.
It is a satisfying story, because it explains something anyone who tracks AI answers has seen with their own eyes: comparison answers do not look like the result list, and nobody outside Google can inspect how they were assembled. So we did the boring thing and tested it. On September 11, 2026 we ran 102 comparison queries against live Google AI Overviews, covering 16 brand pairs, both word orders, two to three repeats each, and stored every answer and every cited source.
The result: order changed the recommendation in zero of 16 pairs. The query-first brand was named first in 49% of runs, which is what a coin flip looks like. What did change the answer, reliably, was how you phrase the query and which Google surface answers it.
The claim, as it spread
The version circulating this month runs roughly like this:
- "Brand A vs Brand B" returns AI Overviews that open with and recommend Brand A about 80% of the time.
- "Brand B vs Brand A" returns the mirror image.
- This holds even when Brand A is far larger and older than Brand B.
- The proposed mechanism: AI Overviews fans your query into sub-searches, and the first brand named in the query anchors those sub-searches.
That last point is what makes the claim credible, and it is also the part that is best documented.
The mechanism is real, which is why the claim travels
Google's guidance for generative AI features, published in May 2026, confirms query fan-out: AI Overviews and AI Mode can expand one query into several related sub-queries, run them against the index in parallel, and synthesize one answer from the combined results. What the sub-queries were is not reported anywhere, Search Console included. So when you see an AI Overview, you are looking at output from a process you cannot watch.
Order effects in language models are real too, and there is peer-reviewed work behind them. A study published in PNAS Nexus in August 2026 tested nine models on resume comparisons and color selection and found position effects that were not mere tie-breakers: in some cases, reordering the options reversed which one the model preferred, and the model then chose the worse option. Work from Columbia Business School found a related pattern in the opposite direction, with ChatGPT picking the first option in a list about 63% of the time across 5,447 prompts, a result that reappeared at 64.29% across 64,800 trials of a simpler setup.
Read those two findings together and the viral claim looks like a natural next step: fan-out exists, LLMs have order quirks, therefore order must decide brand recommendations.
The gap is in the last step. Those experiments ask a model to judge a list of options and pick one. A comparison query in AI Overviews is not that task. It is a synthesis request answered by a retrieval system that grounds the answer in third-party pages. Which is exactly what made the claim worth testing rather than repeating.
What we ran
Queries | 102 against Google AI Overviews, plus 18 against AI Mode |
Brand pairs | 16 (10 software, 6 consumer) |
Word orders | Both, with 2–3 repeats per order |
Phrasing | "X vs Y" for all pairs; also "which is better, X or Y" for 3 pairs |
Surface | Google Search, United States, English, desktop |
Date | September 11, 2026 (single day) |
Capture | SERP API responses, one query per request, storing every AI Overview text element and the full reference list |
Cost | About $0.43 in API credits |
Two design choices matter. First, we ran each query two or three times, because a single flipped answer proves nothing about a system with any randomness in it. Second, we recorded how the answer was built, not just what it said: which brand is named first, how often each is mentioned, what the cited pages look like, and how many citations point at each brand's own domain.
The 16 pairs were chosen to span two situations: software comparisons where both brands are entities people research (Notion vs Asana, Slack vs Microsoft Teams, Semrush vs Ahrefs, Figma vs Sketch) and consumer comparisons where the bigger brand is obvious (Nike vs Adidas, Coca-Cola vs Pepsi, iPhone vs Samsung Galaxy, Netflix vs Disney+, Toyota vs Honda, iRobot Roomba vs Roborock).
Result 1: the answer did not follow the order
Across 100 runs that produced a clean, measurable first mention, the query-first brand was named first 49 times: 49%.
Pair | First named in "A vs B" | First named in "B vs A" | Order changed it? |
|---|---|---|---|
Notion vs Asana | Asana | Asana | No |
Slack vs Microsoft Teams | Microsoft Teams | Microsoft Teams | No |
Wix vs Squarespace | Wix | Wix | No |
Canva vs Adobe Express | Canva | Canva | No |
Mailchimp vs Klaviyo | Klaviyo | Klaviyo | No |
Airtable vs Monday.com | Monday.com | Monday.com | No |
Shopify vs WooCommerce | Shopify | Shopify | No |
Semrush vs Ahrefs | Semrush | Semrush | No |
Figma vs Sketch | Figma | Figma | No |
Zoom vs Google Meet | Google Meet | Google Meet | No |
Nike vs Adidas | Nike | Nike | No |
Coca-Cola vs Pepsi | Coca-Cola | Coca-Cola | No |
iPhone vs Samsung Galaxy | Samsung Galaxy | Samsung Galaxy | No |
Netflix vs Disney+ | Netflix | Netflix | No |
Toyota vs Honda | Toyota | Toyota | No |
iRobot Roomba vs Roborock | Roborock | Roborock | No |
Not one majority flipped. Where an answer had a preference, it kept it in both orders: AI Overviews opened on Asana for "Notion vs Asana" and for "Asana vs Notion," on Monday.com in both Airtable comparisons, on Google Meet in both Zoom comparisons. Nike led both orders against Adidas. Samsung Galaxy led both orders against iPhone.
Repeat runs agreed with each other 93 times out of 100. Six of the seven exceptions came from a single pair under a different phrasing, which turns out to be the actual story. The seventh was one Slack vs Microsoft Teams run that opened on Slack instead of Teams; the next two runs went back to Teams.
One thing this measurement cannot settle: leading an answer is not the same as being recommended. So we also counted recommendation language. Across all 60 software-pair runs, sentences that used win-type words (wins, better choice, recommend, stronger, go-to) split 40 to 39 between the query-first and query-second brand. A dead heat.

Result 2: the two answers were usually the same answer
If order were anchoring the sub-searches, the two versions of a query should retrieve different pages, and the answers should diverge. They mostly did not.
For 8 of the 10 software pairs, the AI Overview for "A vs B" and the one for "B vs A" were word-for-word identical. Compare the first lines of Semrush vs Ahrefs in both orders:
Semrush is an all-in-one marketing platform, while Ahrefs is a specialized SEO and backlink analysis tool.
Same sentence, same order of facts, same closing "choose Semrush if… / choose Ahrefs if…" split, whichever brand the query started with. Only two pairs showed real text drift: Slack vs Microsoft Teams and Shopify vs WooCommerce, and in Shopify's case the drift was confined to a differently worded opening paragraph rather than a different conclusion.
The supporting evidence points the same way. Across 806 reference entries in the software-pair runs, citations to the two brands' own domains came out 36 to 36. Reference page titles that name both brands put the query-first brand first 392 times and the query-second brand first 396 times: 49.7%, another coin flip. Total brand mentions split 337 to 336, or 50.1% to 49.9%.
A system whose retrieval was being steered by whichever brand appeared first in the query would not produce a 36-to-36 citation split and byte-identical prose.
What actually moved the answer
Two things did, and both are more useful than order.
Phrasing. For three pairs we ran the "vs" query and the "which is better, X or Y" query in both orders. The AI Overviews behaved as if these were different questions, and did so identically in both orders.
For Shopify and WooCommerce, "Shopify vs WooCommerce" opened with a product description: Shopify is an all-in-one hosted store builder, and the answer walked through differences. "Which is better, Shopify or WooCommerce" opened with a refusal to rank them: neither is objectively better, it depends on your needs. Same four runs, same pair, same day, opposite first impressions, driven entirely by phrasing rather than position.
For Semrush and Ahrefs, the "which is better" answers contained explicit verdict sentences, and every one of them credited Ahrefs: five of five runs, whether the query read "Semrush or Ahrefs" or "Ahrefs or Semrush." The order was irrelevant next to the fact that the underlying comparison corpus treats Ahrefs as the stronger backlink tool.

Surface. We ran 18 of the same queries through Google AI Mode, the conversational surface. Its behavior differs from AI Overviews in a way that looks like order sensitivity but is not. For Semrush vs Ahrefs, AI Mode opened with "Both Semrush and Ahrefs are…" or "Both Ahrefs and Semrush are…" depending on the query order, mirroring it. And notice which variant it landed on: "Both X and Y," a construction with no ranking implication at all. The name that leads the sentence is a stylistic echo of the prompt.
Underneath that echo, AI Mode's substance was far less stable than AI Overviews: run-to-run text similarity within the same order was only 0.30 to 0.36, and 0.38 across orders, where AI Overviews produced identical text for most pairs. For two of the three pairs, AI Mode's answers were byte-identical across orders, and in the Shopify case it fixed on WooCommerce in its opening both times while AI Overviews fixed on Shopify both times. Each surface has settled phrasing habits, and they are not the same habits.
The order story is not pure fiction
One measurement did move with word order, and it is the one most people are actually watching: the classic result list.
For 5 of the 10 software pairs, the top three traditional organic results differed between "A vs B" and "B vs A," including the pair where the AI Overview was identical in both orders. Reddit ranked near the top in most software comparisons, and which specific thread appeared shifted with phrasing. So there is real order sensitivity in Google Search for comparison queries. It sits in the layer underneath the AI answer, not in the recommendation itself. That distinction is the whole reason this test was worth running, and it is not the finding we expected to write up.
This matters for reporting. If your rank tracker shows your position bouncing between branded and non-branded comparison queries, you are seeing a real effect. It is just not the effect the viral claim describes, and optimizing for it will not move what the AI Overview says about you.
The lever the data does point at
The clearest signal in our runs was not about query construction at all. It was about the comparison corpus each answer was built from.
In Airtable vs Monday.com, the cited sources included 18 Monday.com-owned pages and zero Airtable-owned pages, in both orders. In Mailchimp vs Klaviyo, six Mailchimp-owned citations, zero Klaviyo. Shopify vs WooCommerce produced three WooCommerce-owned citations and no Shopify-owned ones. Meanwhile, on plenty of pairs neither brand's own domain appeared at all: the answer was assembled from third-party comparisons, review sites, and forum threads.
That pattern is consistent with how retrieval actually works. The model does not care which name you typed first; it cares which pages exist and how they describe each brand. Where a brand owns its comparison content, those pages get pulled in. Where it does not, they are missing from the answers of everyone who asks.
How to run your own order-flip test
Fifteen minutes and no API access required:
- Pick five comparison queries where your brand is one of the two names, and where you have a real competitor.
- Run each in both orders, and run each order twice. Four answers per pair, ten pairs of answers total.
- Save the answers before you read them. Screenshot or copy the full text and the cited sources.
- Compare two things separately: which brand is named first, and what the answer says about each brand. Where we found the interesting signal, order was stable and substance was not.
- Watch the citation list, not just the prose. If the competitor's own domain keeps appearing and yours does not, you have found the actual bottleneck.
The AI Overview Citation Checker will show you which pages an AI Overview cites for a query, which is the fastest way to see whether your domain is in the pool at all.
Do this at least twice, on different days, before you conclude anything. Our own run-to-run agreement was 93%, not 100%, and a single flipped run is exactly the kind of noise that turns into a confident blog post.
What to do with this
Do not restructure your comparison pages around word order. There is no evidence that it changes what AI Overviews recommend, and if Google ever did weight it, the fix would be a one-word edit that tells you nothing about your real position.
Do treat phrasing as a real variable. "X vs Y" and "which is better, X or Y" are different queries producing different answers. If your buyers phrase comparisons as questions, you need content that answers the question, not just a spec-by-spec table.
Do check whether your brand's own pages show up in the citations for comparisons you care about. The 18-to-0 split we saw is the most actionable number in this dataset, and it is about publishing, not prompt wording.
Do keep a first-mention metric if you already have one, and treat it as a description of what the answer says, not as a lever you can pull. In our data it was a coin flip that followed the corpus, not the query.
Limits
This is one day of data from one location and language, on desktop, captured through a SERP API rather than a browser session, across 16 brand pairs. Search results vary by personalization, time, and device, and a single day cannot rule out a future change.
We measured which brand is named first, how often it is mentioned, and which sources are cited. None of those is a direct measure of how persuasive the answer is to a reader, and an answer that names your brand first can still leave the impression that the other one wins.
The AI Mode sample was 18 queries across three pairs, enough to see the surface behave differently and not enough to quantify it. Treat that section as a direction, not a rate.
Finally, the mechanism claims in the original story remain unverifiable from outside. Google confirms fan-out happens and does not expose the sub-queries, so nobody can show which sub-queries your query produced. Our test can show that flipping the order did not change the output. It cannot show why.
FAQ
Does word order in a Google query affect AI Overviews? In our test, no. Across 102 AI Overview queries covering 16 brand pairs in both word orders, the query-first brand was named first 49% of the time, and no pair's answer changed direction when we reversed the query. Eight of ten software pairs returned word-for-word identical answers in both orders.
Why do people report order affecting AI recommendations? Because order effects are documented in language models asked to pick from a list, where the first option wins about 63% of the time in published research, and because Google's query fan-out makes the retrieval process invisible. Both facts are real. Neither one means a comparison query in AI Overviews will flip brands when you reorder the names.
What actually changes an AI Overview for a comparison query? Phrasing and surface. "Shopify vs WooCommerce" produced a product-led answer in both orders, while "which is better, Shopify or WooCommerce" produced a "neither is objectively better" answer in both orders. The same queries on AI Mode came back with different framing again, and far more run-to-run variation.
Is it worth optimizing for first mention in AI Overviews? Watch it, do not optimize for it. First mention tracked the source corpus in our data rather than anything we controlled from the query. The controllable variable is whether your own comparison pages are in the citation pool at all: one pair we tested drew 18 competitor-owned citations and zero from the other brand.
Related reading
- How to optimize comparison pages for AI search citations
- Why your "best X" article makes AI recommend competitors
- Which query types AI Overviews actually change, and what they cost you
Author: Ethan Marlowe, GEO Measurement Lead Across 500+ Prompts at Auspia. Ethan writes about AI visibility measurement, prompt-set design, and how to run tests that survive a second look.




