The likely problem
Your pages rank in Google. Maybe top five. But when someone asks ChatGPT or Perplexity "who is the best [your thing] in [your city]," the answer names your competitors, and you don't appear in a single citation.
That's the gap generative engine optimization (GEO) exists to close. GEO is the practice of making a site the kind of source AI answer engines actually cite: extractable structure, clear entities, named facts, third-party signals. And the first step isn't guessing why AI skips you. It's running a citation-gap audit, which is exactly the kind of analysis DeepSeek Harness does well.
This article walks through that audit with a real run, live from a local harness session, and the four fixes it produced.
Symptoms to check first
You're losing AI citations, not rankings. Before the audit, confirm the symptoms:
- [ ] The site ranks on page 1 for its money queries in Google
- [ ] AI answers for those queries name 2-4 brands consistently
- [ ] Your brand is never one of them, even when the site is a better match
- [ ] The same answer surfaces across ChatGPT, Perplexity, and Google AI Overviews
If those four hold, you have a citation gap, and it won't fix itself with more on-page SEO alone.
Why AI engines skip you
Two hard truths from the live audit you're about to read:
1. AI citations are corpus-driven, not web-driven. When ChatGPT answers, it synthesizes text it was trained on, not your live site. If third parties (Reddit threads, Eater and Sprudge listicles, local press) never mention you, the model has nothing to cite. Your pages are invisible to it no matter how well they rank.
2. Being a great site isn't the same as being a retrievable entity. Models bind names to places and categories through repeated co-occurrence: "Moonlight" + "Portland" + "roaster," repeated across listings, reviews, and press. No consistent record, no binding, no citation.

The pipeline that decides whether you get cited: corpus, co-occurrence, answer. Your site enters it only through third-party mentions.
Fix 1 — Run a citation-gap audit in DeepSeek Harness
This is the diagnostic. Paste a real AI answer, describe your site, and get a gap analysis. The prompt (from the prompt library article in this series):
Here is what [ChatGPT/Perplexity/Gemini] answered for "[query]":
"[paste the AI answer verbatim]".
Our site is [site description, e.g., "moonlightroasters.example, a Portland
coffee roaster"]. We are not cited. Analyze the citation gap: why we are
likely not cited, what the cited pages have in common, and 3 concrete fixes.
Be specific.Run it in the web UI composer, or headless for a clean transcript:
dsh --profile headless "PASTE THE PROMPT ABOVE"The real output
This is the actual audit output from a live session. The query was "best specialty coffee roaster Portland"; the AI answer named Tandem Coffee Roasters, Coava Coffee, and Heart Roasters. Our fictional site, moonlightroasters.example, sells whole-bean coffee and subscriptions in Portland, and was not cited.
Why we're likely not cited. LLM citations are corpus-driven, not web-driven. ChatGPT's answer is synthesized from text it was trained on — the brands that most frequently co-occur with "Portland + specialty coffee + single-origin + cafe culture" across Reddit threads, Eater/Sprudge listicles, and local press. It doesn't query live search results or crawl your site at answer time; your own pages are invisible to it unless third parties mention you.
>
The tell is in the answer itself: Tandem Coffee Roasters is in Portland, Maine, not Oregon. This is a text-corpus synthesis (a brand heavily recommended in "best roaster" threads), not a local business lookup. If Moonlight doesn't appear in those threads, it doesn't exist in the model's association graph.
>
Your site's specific gaps: e-commerce-only positioning with no "cafe culture" association · no third-party coverage, so zero training-data citations · weak entity signals, no consistent record tying "Moonlight" to "Portland" + "roaster."
What the cited pages have in common:
>
1. Physical cafe presence and public events (tastings, cupping, bar culture) 2. Transparent single-origin sourcing narratives — farm, producer, process stories on the site 3. Deep third-party editorial coverage — Eater, Sprudge, local papers, r/coffee — the actual source of the citation 4. Consistent name-city-category co-occurrence across every touchpoint
3 concrete fixes for moonlightroasters.example:
>
1. Win third-party mentions — the only thing a model can cite. Pitch "best coffee roaster" roundups, get samples to local writers, run public cupping events and announce them to r/coffee. Add a Press page that links every mention. 2. Mirror the query vocabulary and add schema on-page. Put "specialty coffee roaster in Portland" and "single-origin coffee Portland" in title tags and headings; publish an origin story per coffee; add LocalBusiness/Organization and FAQ schema with consistent NAP, so crawlers bind Moonlight ↔ Portland ↔ roaster as one entity. 3. Close the cafe-culture gap. The cited trio all have cafes and events. Open or partner for a tasting room, host weekly public cuppings, list them on local event calendars.
The short version, straight from the run: you aren't cited because no third party ever wrote about you in a way the model saw. Fix the entity, fix the vocabulary, and earn the coverage — citations follow coverage, not site quality.
Three things to notice about that output. It caught a factual contradiction in the AI answer itself (Maine vs. Oregon). It separated what your site can change (vocabulary, schema) from what it can't (training corpora). And every fix came with a concrete action. That's the quality bar for a good audit.
Verify. The fixes should survive one sanity check: does each one increase the chance a third party writes about you, or the chance a crawler binds you to your category? If a proposed fix does neither, drop it.
Fix 2 — Rewrite the answer-adjacent section
The audit tells you what to say; this step makes your page say it in the structure AI engines quote. Pick the query you want to win, find the section of your page that should answer it, and run this prompt:
Read [file path]. Rewrite the section most likely to be quoted for the query
"[query]" so it matches how AI engines assemble answers: one-sentence
definition up top, named entities, numbers, and no buried lead.The structural rule it follows: an answer engine wants to lift one clean block — definition sentence, named entities, concrete numbers. On the Moonlight site, the "About" page rewrite for "best specialty coffee roaster Portland" would move the mission statement down and put a one-sentence answer up top: "Moonlight Roasters is a specialty coffee roaster in Portland, Oregon, roasting single-origin beans in small batches and shipping subscriptions across the Pacific Northwest." The entities (Portland, single-origin, subscriptions, Pacific Northwest) land in the first two sentences, which is where citation happens.
Verify. Read the rewritten section aloud. If the answer to the query appears in the first two sentences, with a named place and a number or category, it's quotable.
Fix 3 — Make your entity unambiguous
Fix 1's audit runs this check for you, but you can run it standalone on any page:
Read [file path]. List every fact an AI system would need to describe this
business as an entity (name, category, location, offering, pricing,
distinguishing claims), and flag which ones are missing or ambiguous on the
page.Expect the output to be a table of facts, with gaps flagged. Most sites fail two checks immediately: no consistent city-category phrasing in titles and headings, and no pricing (AI answers rarely recommend brands whose prices it can't state). Both are cheap to fix.
Verify. A stranger should be able to read only the first screen of your page and answer: what you are, where you are, what you cost.
Fix 4 — Make it a weekly loop
Citations shift as models update and training corpora change. Turn the audit into a repeatable check (the automation article in this series builds the full scheduled version):
Generate the weekly GEO snapshot for [site] and [3 target queries]: for each
query, paste the current AI answer, note whether we are mentioned or cited,
and list any new sources that appeared. Output as a table. Keep the answer
under 200 words.Run it weekly via headless mode, save the output to a file, and diff against last week. The three things that should alarm you: a citation you had disappeared, a competitor appeared in an answer that previously named neither of you, or a new source surfaced that you could pitch.
When not to do this
GEO work is wasted when the basics are broken. If the site doesn't rank in Google for its own money queries, fix rankings first — AI engines skim the same corpus of coverage and pages that Google's own systems use. And if the business is brand-new with zero third-party footprint, skip the on-page polish and spend the budget on coverage (Fix 1's first fix) before anything else.
FAQ
Does ranking in Google help me get cited by ChatGPT? Indirectly. Both systems lean on the same external coverage and entity signals, so good SEO correlates with citation-friendliness. But ranking alone won't put you in an answer; the corpus has to know you exist.
I can't afford PR. Is GEO pointless? No, but be honest about the ceiling. On-page fixes (vocabulary, schema, structure) win you citations in narrow, factual queries. Broad "best of" answers reward coverage, and the cheapest coverage is often small: local newsletters, niche subreddits, one journalist who covers your category.
Which AI answer engine should I optimize for first? The one your customers actually use. ChatGPT has the largest corpus influence; Perplexity and Google AI Overviews lean more on fresh, retrievable web content, which rewards on-page fixes faster. Run Fix 4 across all three and let the data decide.
Is this the same as traditional SEO? Related but different. SEO optimizes for ranked blue links; GEO optimizes for being named inside an AI-generated answer. This guide's method is the citation-gap audit: paste an answer, analyze the gap, fix entity and coverage. You need both, in that order.
Author: Adrian Cole, GEO Strategy Researcher Analyzing 300+ AI Answer Sets at Auspia. Adrian writes about citation-gap audits, entity optimization, and how AI search engines choose sources.












