How to Find the Sources Shaping AI Answers in Your Industry

Key takeaways

A citation map tells you which pages an AI engine actually pulled from, rather than just whether your brand got mentioned. Here is the full method, plus a runnable setup for Codex, Claude Code, ChatGPT, and Hermes Agent.

What you will finish with

A citation map is a record of which pages an AI engine actually pulled from when it answered your buyers' questions. Not whether your brand appeared in the answer. Which sources the answer was built on.

I have run this on enough categories to know the first result is usually humbling. Teams assume their own site carries the answers. It frequently does not. In one B2B category the vendor's own pages accounted for a rounding error while a review directory nobody had claimed did most of the work.

That is the point of doing this. You stop guessing and start looking at the actual pages.

By the end of this walkthrough you will have:

  • 10 to 20 real buying questions for one product or service category, written the way a customer would ask them
  • Each question run through the AI engines your buyers actually use, with the cited sources recorded and dated
  • Those sources sorted by type, so you can see where your category's answers come from
  • A short list of pages where a competitor is cited and you are not, each with a legitimate next action
  • Owners and a review date, so the map does not rot in a shared drive

Who this is for: an in-house SEO, GEO, or content lead, or an agency strategist who can recommend work even when someone else implements it. No paid tool is required, though one helps at scale.

Prerequisites: a browser, a spreadsheet, and read access to your own site. For the agent versions, a working install of whichever agent you pick, plus a folder you can point it at.

Time: about two hours for the first pass on one category. Roughly 45 minutes of that is running prompts and copying sources. That 45 minutes is the part worth handing to an agent, and the rest is the part that needs your judgment.

Done means: you can name the three or four source types that carry answers in your category, point to specific pages where you are absent and a competitor is present, and hand someone a task with a URL attached. If you only have a list of domains and a vague sense that Reddit matters, you are not done.

Auspia's view: most teams skip straight to "we need more brand mentions" because they never looked at what the answer was actually built from. The map is what turns that vague instinct into a page you can act on.

Before you start: the two decisions that shape everything

Two choices determine what the rest of this produces. Both are easy to get wrong, and both are annoying to undo later.

Pick one category, not your whole business. A company that sells software and also runs a services arm has two different source ecosystems. Map both at once and you get an average that describes neither. Pick the category with the most revenue at stake and leave the rest for next quarter.

Pick your engines deliberately. Google AI Overviews, ChatGPT, Perplexity, and Gemini do not draw from the same pool, and a source that has faded in one may still carry weight in another. Run at least two. If you do not know which two, ask your sales team which surface customers mention, or check your referral traffic for AI domains.

One boundary to set now. This is an observation exercise, not a proof exercise. You are recording what an engine returned on a given day. Answers shift between runs, so one observation is a data point, not a trend. Date everything, or the whole file becomes unusable in a month.

Step 1: Build the question set

Write 10 to 20 questions a customer would ask while choosing. Not keywords. Questions, in the customer's own words.

This is the step people rush, and it is the step that decides whether the rest is worth doing. A question set built from keyword exports produces a map of how people search. A question set built from sales calls produces a map of how people buy. Those are not the same list.

Good sources for the list:

  • Your sales team's most common pre-purchase questions
  • Support tickets from the first 30 days after purchase
  • The "people also ask" block on your main commercial queries
  • Your own keyword research, rewritten into question form
  • Competitor comparison searches your buyers run

Include a mix of question shapes, because they pull different sources:

Question shape

Example

What it tends to surface

Category fit

"What should a mid-size logistics company look for in a routing tool?"

Editorial roundups, review sites

Head-to-head

"Tool A vs Tool B for a 20-person team"

Comparison pages, community threads

Risk and objection

"Is Tool A actually secure enough for healthcare data?"

Forums, compliance writeups, vendor docs

Price and value

"What does Tool A really cost per seat?"

Pricing pages, community threads, review sites

Implementation

"How long does it take to migrate off Tool B?"

Docs, YouTube walkthroughs, forums

Expected output: a spreadsheet with one row per question, a column for question shape, and a column for the engine you will run it in.

Quality check: read the list back and ask whether a real buyer would type each one. If a question only makes sense to someone who already knows your product taxonomy, rewrite it. Ten questions a buyer would actually ask beat fifty that sound like keyword exports.

Recovery path: if you cannot get to ten questions, you have a research problem, not a mapping problem. Pull the last 20 sales-call notes or the last 50 support tickets and extract questions from those before continuing.

Table showing five question shapes - category fit, head-to-head, risk, price, and implementation - and the source types each one tends to surface

Question shape is the hidden variable. A head-to-head question and a risk question about the same product will not return the same sources.

Step 2: Run the questions and capture the sources

This is the mechanical core, and it is where hand-running falls apart. Do five questions by hand so you understand what the output should look like. Then automate the rest.

For each question, in each engine:

  1. Start a fresh conversation. Prior context changes what the engine retrieves.
  2. Ask the question exactly as written.
  3. Open the citation panel or source list. In most engines this is a small icon near the answer, not a visible list.
  4. Record every cited URL, plus the domain and the source type.
  5. Note the date, the engine, and whether the answer cited anything at all.

Record in this shape:

Field

Why it matters

Question ID

Ties every observation back to the buying question

Engine

Lets you separate engine-level patterns later

Date and time

Answers drift; an undated row is unusable in a month

Cited URL

The actual page, not only the domain

Domain

For grouping and frequency counts

Source type

owned, retail/marketplace, community, review directory, editorial

Competitor present

Yes/no, and which one

Our brand present

In the answer, in the sources, both, or neither

That last row is worth dwelling on. A brand can be named in the answer without any of its pages being cited. It can also be cited without being named prominently. Those are two different signals, and if you only track "did we show up," you will misread both. I have watched teams celebrate a mention that came from a page arguing against them.

Expected output: a long table, one row per cited source per question per engine.

Quality check: for at least three rows, click through and confirm the cited page actually supports the claim it was cited for. Engines occasionally cite a page that mentions the topic without answering the question. Those rows are noise and should be marked.

Recovery path: if an engine returns no citations for a question, that is a finding, not a failure. Record it as zero sources and move on. Questions that get answered without any retrieval are questions your own content cannot influence.

Diagram of the capture loop: ask the buying question, open the citation panel, record the cited URL and source type, then tag whether a competitor and your brand appear

The loop is simple and tedious. That combination is exactly what makes it a good candidate for an agent.

Step 3: Sort the sources by type

Now collapse the long table into a picture of your category. Group every cited URL into one of five buckets:

  • Owned: your site, or a competitor's site
  • Retail and marketplace: product listings, price comparison pages, app stores
  • Community and UGC: Reddit, YouTube, Quora, forums, Discord threads that are indexed
  • Review directories and B2B platforms: G2, Capterra, Clutch, Trustpilot, industry-specific directories
  • Independent editorial and reference: trade press, news, Wikipedia, analyst writeups, independent blogs

Count per bucket, per engine, and you have your citation map.

What you are looking for is not a single number. It is the shape. Some categories are carried almost entirely by community threads and trade press, and a brand's own site barely registers. Others are dominated by product listings, because the listing itself contains the answer. The mix follows what buyers ask, not how good anyone's website is. That is a hard thing to accept when you have spent two years on the website.

Expected output: a small table, one row per source type, one column per engine, showing share of cited sources.

Quality check: if one bucket holds more than about 70% of your citations, spot-check five URLs inside it to confirm they are genuinely that type. A single high-traffic aggregator can masquerade as an editorial source while actually being a directory.

Recovery path: if the counts look random, you probably mixed question shapes in a way that hides the pattern. Split the table by question shape and look again. Price questions and risk questions often have completely different source profiles.

Step 4: Find the gaps where a competitor is cited and you are not

Filter the long table to rows where a competitor appears and your brand does not. That filtered list is your worklist. It beats a general "we need more mentions" mandate because every row already has a URL and a customer question attached. You can hand it to someone on Monday and they can start.

For each row, open the cited page and answer three questions:

  1. Does our product genuinely belong on this page?
  2. If yes, what is missing - a listing we never claimed, a comparison we are absent from, a thread where nobody mentioned us, a review we never asked for?
  3. What is the smallest legitimate action that would put us there?

The third question is where discipline matters, because the action is different for every bucket. Getting this wrong is how an SEO ends up cold-pitching a journalist about a Reddit thread.

Source type

Legitimate action

What not to do

Review directory

Claim and complete the profile, ask customers for reviews

Buy reviews or seed fake ones

Community thread

Answer the actual question honestly as a named employee, disclose the affiliation

Astroturf, or drop a link with no contribution

Editorial

Pitch a genuinely useful angle the publication has not covered

Mass-pitch the same press release to 200 domains

Marketplace or retail listing

Fix the specs, images, and description so the facts are retrievable

Keyword-stuff the listing

Owned page

Restructure so the specific fact is easy to find and quote

Add an FAQ block that answers nothing

Expected output: a filtered, prioritized list. Rank by how often the source appears across your question set, not by how prestigious the domain sounds. A forum thread cited on six of your twenty questions outranks a trade publication cited once, and it is also a lot easier to do something about.

Quality check: for each row you keep, confirm you can state the customer question it maps to. If you cannot, drop the row. Pages that get cited for questions nobody in your pipeline asks are a distraction.

Recovery path: if the filtered list is enormous, you have not filtered hard enough. Cap the first pass at ten rows, ordered by citation frequency, and work those before reopening the full list.

Matrix mapping five source types to the legitimate action for each and the tactic to avoid

Same gap, different repair. Treating every missing citation as an outreach problem is how teams end up pitching editors about Reddit threads.

Step 5: Hand the mechanical half to an agent

Steps 2 and 3 burn the most time and require the least judgment. That is the definition of work worth delegating. The question set and the prioritization stay with a person, because those are the two places where a wrong call costs you the whole exercise.

All four agents below can run the same method. What changes is where the method lives and how the output comes back to you. Pick by that, not by which model wins a benchmark this month.

Codex

Best when you want the map to live in a repository and produce a reviewed artifact each run. If your site already lives in git, this is the least friction.

Set up a project folder with a citation-map/ directory, a questions.csv holding your question set, and a SKILL.md that states the method: which engines to check, which fields to record, how to classify a source, and what the output file should look like. Keep one file per run, named by date, so you accumulate history instead of overwriting it.

Ask Codex to run the questions, append rows to the current run file, and produce a summary table of source share by engine. Because the output is a file in a repo, you get a diff you can review before anything is treated as final. That review step is the point: you are checking the classifications, not redoing the capture.

One rule belongs in the skill file no matter what: the agent records what it observed and does not infer a citation it did not see. Invented sources are the most damaging failure mode in this workflow. A written rule plus a spot-check on three rows catches most of it, and the three rows you check should change every run.

Claude Code

Best when the method needs to be read against an explicit written policy, and when you want long-context review of the captured pages.

Put the method in a CLAUDE.md or a project skill, including your source-type definitions and the exact output schema. Point Claude Code at the run folder and ask it to classify each cited URL, then flag any row where the page content does not actually support the citation.

The second job is the one worth paying for. Reading 200 cited pages and judging whether each one genuinely answers the question is miserable for a person and reasonable for an agent with a written standard. Ask for a confidence flag on every row and review only the low-confidence ones yourself.

ChatGPT

Best when you want the lowest setup cost and the work is a single sitting rather than a recurring run.

Paste the method as a custom instruction or a saved project prompt, attach your question list, and work through engines in batches. Ask for the output as a table with the exact columns from Step 2 so it pastes straight into your spreadsheet.

The catch is that ChatGPT is also one of the surfaces you are measuring. If you are mapping ChatGPT's own citations, do that in a clean session with your method unloaded, or you contaminate the observation. Use one session to capture and a different one to organize.

Hermes Agent

Best when you want the method to persist as a skill and improve across runs without you re-explaining it.

Install the citation-map method as a Hermes skill with the question set and output schema, then run it on a cadence - monthly is usually right. Because the skill and its memory persist, corrections you make in month one carry into month two. If the agent misclassified a directory as editorial, fix the rule once.

The tradeoff is that persistent memory needs periodic review. Every few runs, read the skill file and check that the accumulated corrections still describe the method you want. Left alone, it turns into a pile of one-off exceptions that nobody can explain.

Comparison table of Codex, Claude Code, ChatGPT, and Hermes Agent showing what each is best at, how the method is stored, and the main tradeoff

All four run the same five steps. Pick by where your method should live, not by which model scores highest on a benchmark.

Step 6: Verify the map before you act on it

Do not assign work off an unverified map. Five minutes here saves a quarter of misdirected effort.

  • Re-run three questions a week later. If the source mix changes completely, your question set is too volatile or your sample is too small. Record the churn instead of treating the first run as ground truth. Some churn is normal; total turnover is a warning.
  • Confirm the engine, not only the answer. A source that appears in Perplexity may be absent in ChatGPT for the same question. Keep engine columns separate in your summary; a blended total hides exactly the signal you need.
  • Check your own citations too. If your brand is cited, open the page. Sometimes the engine cites a page where you are mentioned critically, or a page you do not control at all. That is useful information, and it is not a win.
  • Confirm the date on every row. An undated citation map cannot be compared to anything later.
  • Sanity-check the source count. If one question returned 40 sources and the rest returned three, check whether you captured a "related" panel rather than the actual citations.

Done means: you can hand a colleague the summary table and one filtered worklist, and they can act without asking you what any row means.

Maintaining the map

Run the full pass quarterly on the same question set, and a light pass monthly on the five questions with the most commercial value. Keep the old runs. The interesting signal is rarely the snapshot. It is the drift: a source type quietly growing, or a competitor who started showing up in a bucket where they used to be absent.

Revisit the question set twice a year. Buying questions change with your product and your market, and a stale question set produces a map of a category you no longer sell into.

One thing to resist: turning this into a dashboard. The output is a worklist with owners and URLs. If nobody has a task at the end of it, the map has not done its job.

FAQ

Do I need a paid AI visibility tool to do this? No. The manual version works for one category and takes an afternoon. Tools earn their cost when you are tracking many categories or markets, or when you want historical trend lines without maintaining the spreadsheet yourself.

How many questions is enough? Ten to twenty per category. Fewer than ten and you cannot tell a pattern from a coincidence. More than twenty on a first pass and you will not finish the analysis, which is the part that produces the worklist.

Should I include questions where we already appear? Yes. Knowing which sources carry you is as useful as knowing which ones miss you. It also tells you what to protect when you rewrite a page.

What if the engine gives a different answer every time I ask? That is normal, and worth recording. Run the same question three times and note the variation. If the source set is completely different each time, treat that question as low-confidence and weight it less when you prioritize.

Can I just use the engine's own source list without opening the pages? For the initial capture, yes. For the rows you plan to act on, no. You have to read the page to judge whether your brand belongs there and what a useful contribution looks like.

How is this different from rank tracking? Rank tracking tells you where your pages appear in a results list. A citation map tells you which pages an answer was assembled from. A page can rank first and never be cited. A page that ranks nowhere can be the backbone of an answer. This is the gap that makes the exercise worth the afternoon.

Author: Ethan Marlowe, GEO Measurement Lead Across 500+ Prompts at Auspia. Ethan writes about prompt tracking, citation reporting, visibility dashboards, and how to tell a real AI visibility change from run-to-run noise.

Explore this topic

Keep following the same growth thread