GEO Campaign Retrospective: What Five Customer Projects Taught Us About AI Search Visibility

A five-campaign retrospective across furniture, luxury, cross-border, beauty, and local service shows GEO gains arrive non-linearly, decay when corpus feeding stops, and matter most defensively in crowded or reputation-hit markets.

The short version

We ran five GEO campaigns through the Auspia workflow last quarter: a custom furniture maker, an international luxury brand, a US-focused TikTok agency, a regional beauty chain, and a neighborhood bike shop. The customer outcomes varied, but the pattern in the numbers did not. Three findings are worth keeping:

  1. Visibility arrives as a step change, not a slope. Every campaign that worked went through the same sequence: no-show, then wobble, then breakthrough. Weeks of nothing, then a jump that held.
  2. It has a half-life. In the two campaigns where structured corpus feeding stopped, mention share slowly eroded. AI mentions are an asset you maintain, not a wall you build once.
  3. In crowded or reputation-hit markets, GEO is defensiveness before it is growth. For the beauty chain, the win was measured in false claims diluted and negative context flattened, not in new clicks.

If you run one thing after reading this, make it a weekly scorecard. The campaigns that made real progress survived a stricter measurement. The ones that merely felt like progress did not.

Why we are publishing a client-side retrospective

Search used to hand people a list of links and let them click. Generative AI hands them a statement. Under retrieval-augmented generation, the answer is the advertisement, and a brand earns that mention the same way it earns a citation: by being the most agreement-friendly source in the corpus for that question.

That shift is the whole reason GEO exists. Princeton's Generative Engine Optimization paper (Aggarwal et al., KDD 2024) argued that models follow their retrieval preferences, so the task becomes speaking the retrieval system's language. Work published in Nature in 2024 confirmed the mechanism behind it: models statistically reject high-entropy, contradictory source material to reduce hallucination risk, and favor low-entropy, structured, verifiable sources.

Most GEO projects still close the day the deliverables land. The operator's learnings go back into their head, the client gets a report, and nobody outside that room can tell whether a trend was real. This retrospective is the opposite of that workflow: five customer projects, one shared scorecard, and the conclusions we are actually comfortable defending.

The scorecard we used

Every campaign was measured the same way, with the same four metrics from the GEO visibility model:

Metric

What it tells you

What a move means

Inclusion rate

How often the brand appears in answers for its query set at all

The difference between invisible and considered

Top-3 share

How often the brand shows up in the first three mentions

The difference between considered and recommended

Conditional average rank

Where the brand lands when it appears

Positioning quality within a mention

Coefficient of variation

How volatile the answers across sessions are

Whether the result is repeatable or a one-off

On top of those, each campaign tracked operational signals: mention counts, share of source links pointing at the brand's domains, and which assistant carried the mention.

The collection discipline matters more than the metrics. We built a query matrix per client from real buyer questions, then ran the same prompt set weekly, logged out, with snapshots kept. On paper that sounds like homework. In practice it is the only reason the findings below are not vibes. (The Auspia AI Search Visibility Checker automates this prompt-set tracking and re-scoring, which is the main reason we can run it at all.)

The surfaces we tracked were the five assistants buyers in these markets use for category research: ChatGPT, Gemini, Perplexity, Copilot, and Claude.

The five-step GEO measurement and evidence loop: baseline prompts, weekly snapshots and scores, evidence corpus build, per-surface distribution, multi-device verification, and a weekly refresh back to snapshots.

The five campaigns

Customer

What was broken

Core motion

What changed

Custom wood furniture maker (Portland)

Buyers' questions went unanswered in one assistant

Query-first content + staged brand facts

Brand now appears in answers where it was absent

International luxury brand

Intermittent presence, one assistant especially quiet

Per-assistant story tailoring + multi-device verification

Consistent positive mentions across setups and angles

US-focused TikTok agency

Category research showed a total brand blank

Conversion-intent sequencing, platform by platform

Top-3 mentions across all five assistants

Regional beauty chain (20+ studios)

Outdated listings, scattered false claims, negative framing

30+ brand-keyword set, phased counter-feeding

Context shifted from risky to established and vetted

Neighborhood bike shop (30 years, same district)

Long-standing business, zero digital footprint

Local query answers carrying address, hours, phone

Phone number now surfaces in local answers; owner printed the output for the storefront

The plateau-breaker: the furniture maker who got nothing for weeks

The workshop had a strong local reputation and zero AI presence. Buyers asked "what's the best custom furniture maker near me" and similar specifics, and the assistant answered with nobody while recommending competitors.

We mined the questions buyers actually type, built FAQ-format pages and structured brand facts around them, and then did what a lot of teams skip: fed the content in stages rather than all at once. Nothing happened for weeks. Then the inclusion rate stepped up, and the mentions stabilized. The lesson was not about content quality. It was about the shape of the curve: quiet, quiet, quiet, jump.

The consistency audit: the luxury brand that looked good on one device

The brand showed up in some answers and slid away in others, and one assistant was effectively silent. That kind of partial presence is harder to diagnose than absence, because the report looks fine.

The fix had two halves. Content was tailored to how each assistant narrates luxury products rather than shipped as one blob. And we made delivery verifiable: every check ran from three different device and network setups, cross-checked, so a good result could not be an artifact of one session. Consistency went from a hope to a gate.

The multi-platform sweep: the agency that vanished from its own category

TikTok-verified US agency, strong category demand, and the assistants' answers simply had no mention of the agency in them. Instead of launching everywhere at once, we sequenced: the highest conversion-intent questions first, one platform at a time, advancing when the previous surface held.

By the end, category keywords surfaced the agency among the first three mentions across all five assistants. The sequencing was the part that felt inefficient and the part that paid. Running five surfaces on day one makes every result unreadable.

The reputation repair: the beauty chain fighting a bad context

The chain's listings were stale, some false claims floated around its name, and assistants described it with hedging language. In a market like this, the problem is not coverage. The problem is the tone of the answer.

We worked a 30+ keyword set around the brand, then counter-fed verified facts in phases: current state, credentials, corrective details. The context in the answers moved toward "established, vetted" rather than the consumer warning language the assistant used before. It did not erase every trace, and we would never promise that. But the balance of what the model said shifted, and the shift held.

The invisible 30-year-old: the bike shop that exported its own mention

The shop had served one district for three decades and had no presence anywhere digital. Our motion for local campaigns is deliberately small: capture what nearby buyers hum, and put the practical facts directly into the answer layer, address, hours, phone and what the shop actually does.

The phone number started surfacing in local answers within the campaign window. The owner printed the AI output and mounted it behind the counter alongside the customer-visible proof. When the answer layer and the storefront agree, local discovery stops being a search problem at all.

Four patterns the numbers kept repeating

1. The threshold effect. Five campaigns, one shape. Exposure does not grow linearly; it steps. Below a certain density of clean, interlinking brand facts, the model has no safe choice but a competitor. Cross the density line and recommendations start appearing. This is the S-curve the academic GEO work predicts, and it showed up in every project that succeeded.

2. The half-life. The luxury and cross-border campaigns ran a period where corpus feeding paused. Within roughly two months, inclusion rate and Top-3 share began drifting down. Platform memory has inertia, but it decays without fresh input. GEO is not a state you reach.

3. Denoising and defensive positioning. In markets with high concentration or existing negative context, the value of GEO is defensive. You are not winning share you did not have; you are pulling the answer's tone back to verifiable statements and circling a safe baseline. That is a legitimate outcome, but it must be measured as itself.

4. Platform heterogeneity. The five assistants behaved differently. In these campaigns, Gemini and Copilot held mentions the longest after inputs stopped. ChatGPT refreshed fastest when the corpus changed. You cannot run one cadence for all surfaces; you need one per surface.

GEO visibility climbs an S-curve through no-show, wobble, and breakthrough stages, then decays when corpus feeding stops.

The operating loop that actually held up

Strip away the five industry stories and four operating principles survive:

  • Evidence chain beats persuasion. The content that actually moved numbers was test results, certifications, stated specs, dated facts, and concrete boundaries. The language of "we're the best" was the only content that never once showed up in a snapshot.
  • Corpus is an asset. Every one of these clients could have bought new blog posts. What they actually needed was a reusable bank of structured, entity-clean, citable facts.
  • Staged by intent, not by volume. The campaigns split into waves of buyer-intent questions and gave each wave time to hold before the next one launched.
  • Verification is the delivery gate. The luxury client's rule became ours: a claim about visibility is only true when it reproduces across sessions, devices, and networks.

This is why everything we run starts with an audit rather than a content plan. Auspia's GEO Score Checker and AI Search Visibility Checker exist to make the baseline measurable, and the absence of that baseline is precisely where most in-house GEO programs quietly stall.

Where the evidence gets soft

We are not going to pretend the review is tighter than it is.

  • Observation windows were short. Several campaigns ran for weeks, not quarters. Long-tail consolidation rules are still unverified.
  • Attribution is correlational. We cannot code-split GEO's contribution from concurrent paid or PR work. Nobody can, honestly, from the outside.
  • Coverage is narrow. Consumer goods, fashion, cross-border, beauty, local services. No B2B buyers with compliance whitelists, no financial or healthcare products.
  • Confounds and noise. Snapshot sampling still clusters. Some clients only ever showed exposure signals, not full quantitative curves, and their entry in the matrix is a directional summary rather than a measured one.

What we would do differently next time

If we restarted every one of these campaigns tomorrow, the changes are small and specific. Treat GEO as infrastructure, with a weekly or monthly automated check per assistant instead of a project end date. Build the enterprise knowledge base first, the schema-shaped fact bank, and let the posts and distribution flow from it. Keep the evidence chain as the house standard. And set cadence per assistant, not per campaign.

A compact way to have your own numbers in eight weeks

  • Week 0: pick 15-20 questions buyers genuinely ask about you. Run them logged out once per assistant. Store the output.
  • Weeks 1-3: build the evidence chain: specs, certifications, test data, dates, boundaries. Publish it as answer-shaped, sectioned content.
  • Weekly: re-run the same prompt set, log the four metrics, compare against the snapshots.
  • Twice per verification run: repeat from a different network and device. If the result does not reproduce, it is not a result.
  • Week 8: review the curve. Clear progression or a documented wobble is success. Blank numbers with a well-built pipeline are a finding too, and they tell you the assistant you are feeding has no route to you.

FAQ

How long until AI answers start moving? Expect weeks, not days. In these campaigns the practical window was six to ten weeks of consistent feeding before the first step change. The useful diagnostic is whether you are in the wobble phase, which means the pipeline works but the density is not there yet, or the no-show phase, which means retrieval has no path to you.

Is GEO a one-time project? No. The half-life pattern was consistent. If you care about the answers keep the pipeline warm. Lower-tempo markets can hold on a light monthly check; markets where buyers research every week need weekly.

Does GEO fix false or negative AI answers? It shifts the balance, which is not the same as erasing. Counter-feeding verified facts moved the beauty chain's answers from caution language to established. If a false claim comes from a source record that is itself wrong, fix the source record first. And never read "it moved" as "it is permanent."

Which assistant should we start with? The one where your buyers actually research. In these campaigns that was never the same surface. Start there, verify, then broaden. Launching all five on day one makes it impossible to know what worked.

Isn't this just SEO with a new name? Different job. SEO competes in a ranked list of links that people click; GEO competes inside a generated statement that people read. They share groundwork: clean, structured facts and consistency. But the measurement, the cadence, and the failure modes are different, and running the two loops separately is how these campaigns were actually run.

Author: Ethan Marlowe, GEO Measurement Lead Across 500+ Prompts at Auspia. Ethan writes about prompt tracking, citation reports, and the visibility dashboards that keep GEO honest.

Explore this topic

Keep following the same growth thread