How to Build a GEO Visibility Bot in Grok Bot

Key takeaways

A working setup for a Grok Bot that checks whether ChatGPT, Perplexity, Gemini, and AI Overviews cite your brand across a fixed set of buyer questions, and logs the answer every week.

This guide builds one Grok Bot that answers a narrow, useful question every week: when a buyer asks one of your tracked questions, does an AI answer mention your brand, and does it cite you as a source?

Who this is for: a growth or SEO lead who already tracks rankings and now needs to track AI answer visibility.

What you finish with: a GEO Bot with a fixed question set, a weekly routine that checks four AI surfaces, and a citation log you can read as a trend.

Before you start: a Grok Bot with access, a list of 20 to 40 real buyer questions, and a dedicated account for any platform that requires a login. Budget about an hour for the first build.

Done means: you have one week of logged results, you have hand-verified at least three of them against the live answer, and you know which surfaces your Bot can and cannot read reliably.

Why a fixed question set matters more than a big one

The instinct is to track everything. Resist it. AI answers are non-deterministic, so a changing question set produces a changing baseline, and a changing baseline produces a chart that means nothing.

Pick 20 to 40 questions that a real buyer would type, group them by intent, and freeze the list. Change it once a quarter, and note the change when you do.

A useful grouping:

Intent group

Example question shape

What a mention tells you

Category

"What is the best tool for ___?"

Whether you are in the consideration set at all

Comparison

"_ vs _"

Whether you appear in head-to-head answers

Problem

"How do I fix ___?"

Whether you are cited as a source, not just a product

Brand

"Is _ legit?" or "What is _?"

Whether the answer about you is accurate

The brand group is the one teams skip and then regret. If an AI answer describes your product incorrectly, that is a visibility problem whether or not you rank.

The official Grok Bot use cases page listing the work categories a Bot can be assigned

Grok Bot is organized around job categories rather than a single chat. Source: x.ai/bot/use-cases, captured September 23, 2026.

Step 1: Create the GEO Bot and give it the question set

Create a new Bot named GEO Visibility. Do not reuse your technical auditor. Separate memory, separate job.

Paste your frozen question list into the Bot and save it as a file on the shared machine so every run reads the same source:

text
Create a file at /geo/questions.md on the shared machine.

Write this exact list into it, one question per line, grouped by intent.
Do not rephrase, shorten, or add questions.

[CATEGORY]
What is the best AI visibility tool for B2B SaaS?
...

[COMPARISON]
...

[PROBLEM]
...

[BRAND]
...

Confirm the file exists and report the total line count.

Expected output: the Bot confirms the file path and the line count matches your list.

Quality check: the line count is exactly what you pasted. If it is higher, the Bot added questions.

Recovery: if the count is wrong, ask the Bot to show the file contents and correct it before moving on. A question set that drifts on day one will drift every week.

Step 2: Write the monitoring skill

Save this as a skill. The critical sections are the proof rule and the boundaries. Without them you get a confident summary you cannot audit.

text
ROLE
GEO visibility monitor.

INPUTS
- Question set: /geo/questions.md
- Surfaces to check: ChatGPT, Perplexity, Gemini, Google AI Overviews
- Brand name: [YOUR BRAND]
- Brand domains: [yourdomain.com]

WHAT TO RECORD FOR EACH QUESTION AND SURFACE
1. Did the answer mention the brand? yes / no
2. Did the answer cite a brand-owned URL? yes / no
3. Which sources were cited? list the domains
4. The exact sentence where the brand appears, if any

PROOF
- Quote the exact sentence for every mention.
- Record the source domains exactly as shown.
- If you cannot see the answer, record "not accessible" rather than guessing.
- Never infer a mention from a citation alone, or a citation from a mention.

OUTPUT
Append one row per question and surface to /geo/log-{YYYY-MM-DD}.csv:
date, question, intent_group, surface, mentioned, cited, source_domains, quote

Then write a short summary to /geo/summary-{YYYY-MM-DD}.md:
- total mentions by surface
- total citations by surface
- questions where the brand appears in no surface
- any answer that describes the brand incorrectly

BOUNDARIES
- Read-only. Do not post, comment, or submit anything.
- Do not log in to any platform you were not explicitly given.
- If a surface requires a login you do not have, record "not accessible" and continue.
- Never fabricate a result to fill a row.

Expected output: the skill is saved with all four sections.

Quality check: the proof rule explicitly forbids inferring a mention from a citation. That distinction is the single most common source of bad GEO reporting.

Recovery: if the Bot starts filling rows with plausible-looking quotes you cannot find, the proof rule is not being enforced. Ask it to show the live page for one row before continuing.

The official Grok Bot documentation on skills and routines, where the monitoring procedure is saved

The monitoring procedure is saved as a skill so every weekly run follows the same steps. Source: docs.x.ai/grok-bot/skills-routines-and-automations, captured September 23, 2026.

Step 3: Run one surface first

Do not run all four surfaces on the first attempt. Run one, verify it, then add the rest.

text
Run the GEO visibility skill for the CATEGORY questions only.
Check Perplexity only.
Write the results to the log file and show me the table.

Expected output: a table with one row per category question, each with a mention flag, a citation flag, source domains, and a quote where relevant.

Quality check: pick two rows and open the same question in Perplexity yourself. The quote the Bot recorded should appear in the live answer.

Recovery: if the quotes do not match, the Bot may be reading a cached or personalized answer. Ask it to reload and re-check one question, and confirm it is not signed into a personalized account.

Step 4: Add the remaining surfaces

Once one surface is verified, add the others one at a time. Each surface behaves differently, and you want to know which one is unreliable before it contaminates your log.

text
Now run the same CATEGORY questions across ChatGPT, Gemini, and Google AI Overviews.
Keep the same output format.
If a surface is not accessible, record "not accessible" and continue.

Expected output: the log now has one row per question per surface.

Quality check: the "not accessible" count is honest. A log with zero inaccessible rows across four surfaces on a first run is suspicious. Most teams hit at least one wall.

Recovery: if the Bot claims access to a surface you never logged into, ask it to show the page it read. Do not accept a summary.

Step 5: Turn it into a weekly routine

text
Save this as a routine called "Weekly GEO visibility check".
Run every Monday at 09:00.
Use the frozen question set at /geo/questions.md.
Write the log and the summary to /geo/.
Do not modify the question set.

Expected output: the routine is listed with a Monday 09:00 schedule.

Quality check: run it once manually. A routine that has never completed a run is not verified.

Recovery: if the scheduled run fails but the manual run works, a session has usually expired. Re-authenticate on the shared machine.

Step 6: Read the log as a trend, not a score

After three or four weeks you will have something more useful than a single number: a trend per surface. Two patterns are worth watching.

Mentions without citations. The AI names your brand but does not cite your site. That usually means your brand is known but your pages are not the source being used. This is a content and entity problem, not a technical one.

Citations without mentions. The AI uses your page as a source but does not name you. That is often fine, but it means your brand name is not strongly associated with the topic in the source material.

The Auspia AI Search Visibility Checker, which gives a first-pass read on how AI surfaces describe a brand

Use the Auspia AI Search Visibility Checker for a quick external read, then let the Bot track the trend over time. Source: auspia.ai/tools/ai-search-visibility-checker, captured September 23, 2026.

If you want a fast baseline before you build the Bot, run the AI Search Visibility Checker first. It gives you a starting point to compare the Bot's log against.

Verification checklist

  • [ ] The question set is frozen in a file, not retyped each run.
  • [ ] The skill forbids inferring mentions from citations.
  • [ ] One surface was verified by hand before adding the others.
  • [ ] At least three logged rows were checked against the live answer.
  • [ ] Inaccessible surfaces are recorded honestly, not skipped silently.
  • [ ] The routine completed at least one scheduled run.
  • [ ] You can explain the difference between a mention and a citation in your own log.

Common mistakes

Tracking too many questions. Twenty to forty is plenty. A hundred questions produces a hundred noisy rows.

Changing the question set mid-quarter. You lose the trend. Freeze it, then change it deliberately and note the date.

Treating a mention as a citation. They are different signals with different fixes. Keep them in separate columns.

Trusting a first run across all four surfaces. Verify one surface, then add the rest. You want to know which one lies.

Reporting the number without the quote. A mention count with no quoted sentence is not auditable. Keep the quotes.

FAQ

Can Grok Bot read AI Overviews reliably? It can read the page, but AI Overviews vary by query, location, and personalization. Treat those rows as directional and always keep the quote so you can check it.

Do I need logins for ChatGPT, Perplexity, and Gemini? Some checks work logged out. Where a surface requires a session, use a dedicated account and record honestly when a surface is inaccessible.

How is this different from a rank tracker? A rank tracker reports position in a list. This reports whether an AI answer mentions and cites you, which is a different signal with a different fix.

How long before the data is useful? Three to four weekly runs give you a trend. A single run only tells you today's state.

Should the Bot also try to improve the answers? No. Keep this Bot read-only. Improvement work belongs in a separate Bot or a human workflow, so your measurement stays clean.

Author: Ethan Marlowe, GEO Measurement Lead Across 500+ Prompts at Auspia. Ethan writes about prompt tracking, citation reporting, and building visibility dashboards that teams actually use.

Explore this topic

Keep following the same growth thread