I have a habit of re-running the same question in ChatGPT, Gemini and Perplexity every few months and screenshotting the answers. It is the cheapest brand audit I know.
In June 2026, the prompt "best water purifier for a family with a newborn" returned five brands. A manufacturer founded more than a hundred years ago, with a patent portfolio larger than most of its category, was not one of them. It did not appear in the answer, in the citations, or in the follow-up questions. When the same assistant was asked about the brand by name, it described the company as "a large water treatment equipment manufacturer". True, useless, and nearly identical to how it described four competitors.
By September 2026 the same prompt showed the brand in the top three recommendations, with a description that mentioned its sterilization technology and the models suited to small apartments. A second prompt, this one about removing residual chlorine without a storage tank, put the brand first.
A Chinese-language write-up published by the brand's growth team in September 2026 documents what happened between those two dates. I read it twice, extracted the method and the reported numbers, and rebuilt both for teams working in English-language markets. What follows is the teardown: what was broken, what they changed, what the numbers say, and which parts I would copy without hesitation.
One caveat before any of the percentages. The results are vendor-reported by the team that ran the project. There was no control group, no third-party measurement panel, and no published sample size for the prompt checks. Read them as directional evidence about a method, not as an industry benchmark.
What was broken was not content volume
The team's first assumption was wrong, and it is the same wrong assumption I run into most weeks: they believed they had a volume problem.
They were publishing. Product pages, blog posts, trade press coverage, a healthy set of reviews. The brand ranked well in classic search. If you had looked only at Google Search Console, you would have concluded that visibility was fine.
The gap showed up in three places once they started asking assistants directly.
The model's framing of the brand was generic. Ask a language model what the company is known for and it reached for the lowest-common-denominator label: a water treatment manufacturer. The two things that actually differentiate the products, sterilization patents and a compact installation design, appeared nowhere. That is a description problem before it is a ranking problem. Recommendation starts from a model having a specific, retrievable reason to name you.
Pain-point prompts went to competitors. Nobody with a newborn types "best brand". They type something closer to "does a reverse osmosis system remove the minerals a baby needs" or "water purifier that does not need installation under the sink". The brand's content answered the questions the marketing team found interesting, not the questions buyers ask when they are stuck between two options.
The content was written for a human reading a page, not for a machine extracting facts. Long narrative paragraphs, adjectives in place of measurements, almost no structured comparison against named alternatives, no FAQ blocks, and no consistent numbers across pages. When a retrieval system tries to summarize what a brand is good at, material like that offers nothing clean to lift. The model is left to average whatever the rest of the web says, which is usually the same generic category description.
None of those three problems is fixed by publishing more articles. That is why the method that follows is worth reading even if you already have a content calendar.
Move one: map intent instead of chasing keywords
The rebuild started with a map, not a spreadsheet of keywords. The team built 15 intent scenarios across three layers, and each layer exists for a different kind of decision.
The first layer is generic brand decision terms. These are the prompts where the buyer has already accepted the category and is choosing a name. "Best water purifier for a newborn", "most reliable water purifier brand". There are not many of them, and they are the most competitive.
The second layer is core pain-point scenarios. These are the prompts where a specific problem drives the search: sensitive skin, chlorine smell, hard water, a rented apartment where drilling into the countertop is not an option. Volume is lower, intent is much higher, and the brand list the model returns is shorter and more volatile.
The third layer is function and model long-tail questions. These are the comparisons at the bottom of the funnel: tankless versus tank, how often a filter needs replacing, whether a specific model fits under a 60cm counter. Assistants get asked these constantly, usually after a user has already seen three brand names.
Mapping the three layers to the buyer's journey changed what the team wrote. Awareness and word-of-mouth stages were left alone on purpose. Exploration and evaluation, where a model is assembling candidate lists and arguing for one of them, got everything.
That scoping decision is the part most teams get wrong. A GEO program that tries to cover all five stages of the journey at once usually produces content that is too general to be extractable at any of them.

Move two: give the model one story it can repeat
The second move was semantic discipline, and it is unglamorous. The team standardized the brand's vocabulary and reduced it to two or three selling-point labels that appear in the same words everywhere: on the site, in press materials, in product descriptions, in answers to journalist questions, in support documentation.
Alongside that, they built three content assets designed for extraction rather than for reading pleasure.
FAQ modules came first. Questions used as headings, one answer per paragraph, numbers and units instead of adjectives, and a short list of the sub-questions each answer resolves. This is the format an assistant can quote without paraphrasing itself into vagueness.
Comparison matrices came second. Their models against named competitors, four or five attributes at a time, each row carrying the source of the figure. Naming competitors feels risky to most marketing teams. In practice it helps, because it gives the model a structured reason to place you in a comparison instead of describing you in a category.
Closed-loop narratives came third, and this is the piece I would steal first. Instead of a product page that lists features, each page follows problem, mechanism, solution, model: the pain point in the buyer's own words, the physical mechanism that addresses it, the product decision that follows from that mechanism, and finally the specific model numbers. It reads like an explainer and it parses like an argument. When the assistant has to answer "which one for a small apartment", the reasoning is already laid out in order.
The rule they enforced across all three assets was blunt: no performance claim without a number, and no superlative without a source. Claims like "industry-leading purification" were rewritten or deleted. In my experience this single rule does more for answer visibility than any formatting trick, because a model that cannot verify a claim about you will mention a competitor whose claim it can verify.
Move three: publish where the models actually read
The third move treated the open web as the medium, not the brand site.
Owned channels came first, because they are the source of truth the rest of the material repeats. Then came distribution: roughly 200 articles over three months, of which 45 were original main pieces, with each main piece adapted into platform-specific versions. Paid placements, contributed pieces, and official accounts were matched to the source preferences of the assistants they cared about.
Their framing of the logic is worth quoting in substance: the model does not decide what you are by reading your website once. It averages what the web says about you, weighted by how often that claim appears in places it considers authoritative. A brand with a perfect website and no off-site corroboration is a brand with a thin evidence base.
I would add a practical check to that. Before writing a single new article, search your own brand name in three assistants and look at what sources they cite. The list of domains that appears is your actual media plan, and it is usually not the list in the marketing deck.
What moved
Here is the reported before-and-after, kept as reported rather than rounded up to sound better.
Prompt layer | Reported visibility before | Reported visibility after | Reported change | Reported top-choice or top-three rate |
|---|---|---|---|---|
Generic brand decision terms | 12% | 38% | +217% | Top three: below 10% to above 35% |
Core pain-point scenarios | 18% | 72% | +300% | First choice: below 20% to above 65% |
Function and model long-tail questions | 15% | 68% | +353% | Combined recommendation: below 10% to above 55% |
Two supporting numbers deserve as much attention as the headline ones. The number of distinct decision attributes the models associated with the brand went from one generic label to four or more. Coverage of long-tail question phrasings went from roughly one in five to essentially all of them.
That second figure is the quiet one. A brand that is described with four attributes can be recommended for four different reasons. A brand with one attribute competes for one slot in every answer, which is a much worse position than a single headline percentage suggests.
Their measurement method is worth stating plainly because it is reproducible without buying anything: a fixed list of prompts, run on a fixed cadence, in the assistants their buyers actually use, recording whether the brand appeared, in what position, and with which description. Answers are logged with dates so that a change can be attributed to something.
Auspia's own AI search visibility checks run on exactly that pattern, and the only real difference at scale is tooling. Manual sampling is fine up to a few dozen prompts. Past that, you need prompt tracking, citation reporting and a way to diff two weeks of answers, or the program quietly becomes a slide deck.
The loop, and the two holes they left open
The operating model behind the results is a four-stage loop: demand research feeds content creation, content creation feeds source distribution, distribution feeds measurement, and measurement goes back into demand research. Each stage has an owner and an output, and the loop is the reason the program kept improving instead of plateauing after the first push.

They also documented what the loop did not cover, which I respect more than the results table.
The buyer journey has five stages: awareness, exploration, evaluation, decision, word of mouth. The rebuild concentrated on exploration and evaluation. Awareness was left to existing brand marketing, and word of mouth was not systematically engineered at all, even though reviews and community threads are a common citation source.
The second hole is modality. All the assets were text. Video content, product demonstrations and their transcripts were not part of the program, which leaves the growing share of assistant and multimodal answer surfaces unaddressed. If your competitors have product walkthroughs with clean transcripts and you have a PDF spec sheet, you are absent from a set of answers you cannot even see.
A third, more commercial hole sits downstream of the answer. Being recommended by an assistant hands a buyer to a product page, a marketplace listing or a store. Product titles, listing copy, review structure and FAQ-style detail sections all decide whether that handoff converts. Optimizing the recommendation without touching the landing surface is how a team ends up with impressive visibility numbers and flat revenue.
What I would copy, and what I would change
I would copy the intent map, especially the decision-stage scoping. It is the cheapest part of the program and it prevents the most expensive mistake, which is writing a large amount of content that answers nothing specific.
I would copy the vocabulary discipline and the closed-loop page structure. Both are one-time setup costs that keep paying out every time an assistant decides how to describe you.
I would copy the source plan, with the honesty it requires: your own site is not the medium.
What I would change is measurement.
Vendor-reported case studies almost never show the counterfactual, because the vendor has no reason to run one. You can get most of it cheaply by selecting five scenarios you deliberately do not optimize for and checking whether they move anyway. If they do, your gains may be coming from general brand momentum rather than from the work.
I would also push back on the volume. Two hundred articles in three months is a real content operation, and most teams cannot sustain it. The same method run on five scenarios, executed properly, will teach you more in a quarter than fifty shallow articles will.
And I would watch the failure mode this approach creates. Semantic consistency can slide into a vocabulary that is stable and empty. If your two or three labels are not anchored to a verifiable product difference, the assistant will keep describing you generically, and it will have good reason to.
A thirty-day version
If you want to start this week rather than next quarter, the compressed version looks like this.
Pick three assistants and write down fifteen prompts across the three layers, using the phrasing buyers actually use, not the phrasing your marketing team prefers. Run them today and save the answers with dates. That document is your baseline, and without it you will never be able to prove anything.
For each prompt, note which brands appear and which domains are cited. The citation list becomes your media plan.
Then rewrite one product page as a closed-loop narrative and add one FAQ block with numbers, units and named comparisons. Run the same fifteen prompts again in two weeks and compare descriptions rather than positions. Descriptions change first, and they are the leading indicator that a model is starting to understand what you actually make.
Keep the five-stage journey map beside the prompt list. Circle the stages you are not covering, and decide out loud whether that is a choice or an oversight.
Why this case matters beyond one brand
Most GEO writing treats the discipline as either mysticism or formatting. This case is neither. It is a structured content engineering project with a research stage, an asset specification, a distribution plan and a measurement cadence, and its results are proportionate to how boring the work was.
The part I keep thinking about is the long-tail coverage number. Moving from one in five question phrasings to all of them is not a trick about headings or schema. It is what happens when a team stops writing the questions it wants to answer and starts answering the ones buyers bring to an assistant at eleven at night, when no salesperson is awake to intercept them.
That is the whole opportunity. The recommendation now happens before anyone visits your site, and it is assembled from evidence the web already holds about you.
Author: Amelia Ross, AI Brand Positioning Strategist at Auspia. Amelia works on how brands are described, categorized and recommended by AI assistants, and on the content systems that change those descriptions.




