Every AI brand monitoring plan starts with a prompt list, and almost every prompt list is built from what the team already believes. Twenty prompts feels thorough. Fifty feels exhaustive. Nobody checks whether the list is still discovering anything.
We ran 20 prompts across three AI surfaces in September 2026 and tracked one thing: where the sources came from. Not mentions of our own brand, but the full set of domains the answers pulled in. If a prompt list is complete, adding prompts should stop finding new sources. It did not.
What we measured
Item | Setting |
|---|---|
Prompts | 20 buyer questions about AI visibility tooling |
Surfaces | ChatGPT (gpt-5), Gemini 2.5 Flash, Perplexity (sonar), web search enabled |
Answers collected | 60 |
Distinct sources found | 432 domains |
Cost of one full pass | $1.444 |
Repeat coverage | 5 prompts run twice, same day |
We treated each of the 60 answers as one observation of the source universe and counted distinct domains, not citation slots. Gemini returned 989 source annotations across its 20 answers, but many of those pointed at the same page, so the domain count is the honest unit here. The annotation-level data is in the AI visibility checker analysis.
Result 1: the list never closes
This is the part that decides how a monitoring program should be built. After each prompt we added its new domains to everything found so far.
Prompts run | Distinct sources found | Share of the 20 prompt total |
|---|---|---|
1 | 34 | 8 percent |
3 | 91 | 21 percent |
5 | 134 | 31 percent |
10 | 235 | 54 percent |
15 | 336 | 78 percent |
20 | 432 | 100 percent |
The last prompt added 14 domains that no earlier prompt had surfaced. The smallest marginal gain across all 20 prompts was 14, and the largest was 34. There is no point in this range where one more prompt stops paying for itself, which means a prompt list is not a sample you complete, it is a sample you keep drawing.
The practical version: a five prompt list built from memory covers roughly a third of what twenty prompts cover. That is the gap between a monitoring program that notices category change and one that reports the same five answers every month.

Result 2: each surface sees a different third of the same space
The 432 domains were not evenly distributed. Each surface reached a different slice of the same question set.
Surface | Distinct sources | Share of the 432 | Median sources per answer |
|---|---|---|---|
ChatGPT | 47 | 11 percent | 0 |
Gemini 2.5 Flash | 184 | 43 percent | 12 |
Perplexity | 285 | 66 percent | 19 |
Gemini and Perplexity between them found 400 of the 432 domains, so ChatGPT added 32 domains nothing else surfaced. On the 12 prompts where ChatGPT did not search, that surface contributed nothing at all, which is why its median is zero.
The consequence for monitoring is uncomfortable. If your tracker watches one surface, you are not running a smaller version of a three surface program. You are reporting on a different source population, and the answer to "are we being cited more this month" may simply be "we changed nothing and the surface did". We measured the same effect at answer level in AI visibility tracker run variance.
Result 3: two surfaces rarely agree on the same prompt
We compared the source lists Gemini and Perplexity produced for the same prompt, on the same day, in the same language.
- The two lists shared a mean of 2.5 domains out of roughly 20 each.
- The mean Jaccard overlap was 0.089.
- On only 3 of the 8 prompts where ChatGPT also returned sources did all three surfaces share even one domain, and the largest three-way overlap was 2 domains.
- Full agreement across all three surfaces on a single prompt never happened.
Two answers to the same question, built from almost disjoint sets of pages. That is the strongest argument against single surface reporting we found, and it is also good news: a new page has many more chances to be picked up than a single source list suggests.
Result 4: most sources appear once
The 432 domains are extremely unevenly distributed.
- 324 of the 432 domains, or 75 percent, appeared in exactly one of the 60 answers.
- Only 17 domains appeared in 5 or more answers.
- The recurring core is small and predictable: ahrefs.com in 16 answers, semrush.com in 15, reddit.com in 14, then frase.io at 9, and a cluster at 6 that includes otterly.ai, tryprofound.com, llmpulse.ai, cognizo.ai, and searchenginejournal.com.
This splits the monitoring job in two, and the two halves want different cadences. The recurring core is the tier you re-check often, because a change there is a real shift in a shared source. The one-off tail is the tier you sample, because 75 percent of the universe is noise until it repeats.

What to monitor instead of a complete list
Layer | What goes in it | Cadence | Monthly data cost |
|---|---|---|---|
Tier 1: recurring core | The 17 domains that repeat, plus your own domain and your three closest competitors | Weekly on one surface | $0.48 on Perplexity, $3.23 on Gemini |
Tier 2: prompt panel | 15 to 20 prompts covering category, comparison, pricing, and problem framing | Monthly across all three surfaces | $1.44 per pass |
Tier 3: prompt discovery | 5 new prompts per month, written from sales calls and support tickets | Quarterly review, then promote or drop | Included in the monthly pass |
Three rules came out of running this.
- Log the prompt class, not just the prompt. Our widest answer came from a prompt about getting cited at all, not from a tool comparison, and the narrowest came from a comparison. Classes behave differently, and the class is what you report on.
- Keep the surface in every record. A citation count without a surface name is not comparable, because the list lengths differ by a factor of three.
- Rotate the discovery prompts. Twenty prompts found 432 domains, and the curve says a different twenty would find a different few hundred. Promotion and retirement is the only mechanism that keeps the panel honest.
Cost is not the constraint here. A full three surface pass of 20 prompts cost $1.444, so a monthly program runs about $17 a year on data, before review time. The constraint is that review time is the expensive half, which is why the tier structure matters more than the list length. Google's own AI surface is not in this panel at all and needs a different instrument, as our AI Overview tracker explains.
Limits
- One day, one location, English, three surfaces, one model version each.
- Domain counts strip duplicates inside an answer, so a page cited five times counts once.
- Prompt classes were assigned by us, not by any external taxonomy.
- The cost figures cover answer generation only.
- ChatGPT's contribution is depressed by the 12 prompts where it did not search, which is a property of that surface rather than of the prompt set.
FAQ
How many prompts do I need for AI brand monitoring?
There is no completion point in our data. Twenty prompts found 432 sources and the twentieth still added 14 new ones. Choose a panel size you can review monthly, then spend the effort on prompt quality and rotation rather than on list length.
Should I monitor one AI surface or several?
Several, and report them separately. The two heaviest surfaces shared a mean of 2.5 sources per prompt, and full three-way agreement never happened. A single surface program measures that surface, not your category.
Is a domain that appears once worth tracking?
Only as an observation. 324 of 432 domains appeared once, so a single appearance is closer to noise than to a signal. Promote a domain to the recurring tier after it shows up in a second answer, ideally on a second surface.
How often should the panel run?
Monthly for the full panel, weekly for the recurring core. The recurring list is short enough that a weekly pass on one cheap surface costs under fifty cents a month, and the monthly pass is where new sources enter.
Auspia view: stop sizing prompt lists by how thorough they feel and start measuring when they stop finding anything. In our data that point never arrived, which means the right unit of planning is prompt classes and rotation, not one longer list. Tier the recurring sources, sample the tail, and keep the surface name attached to every number you report.
Elena Shaw




