The short version
For years, GEO work had an uncomfortable measurement problem: teams could see AI answers, but they could not reliably prove why one source appeared, how often it appeared, or whether visibility was improving.
That is changing. Bing began exposing AI Performance data in Bing Webmaster Tools in February 2026, then added Intents, Topics, Citation Share, and Compare in June. Google launched Search Generative AI performance reports in Search Console on June 3, 2026, with dedicated visibility reporting for AI Overviews, AI Mode, and generative AI features in Discover.
This does not mean the black box is gone. It means the edges of the box are finally visible.
The practical takeaway for growth teams is simple: GEO can no longer be sold or managed as "we wrote AI-friendly content and saw some screenshots." It needs a measurement routine: which pages appear, which AI query contexts retrieve them, which topics are gaining or losing coverage, and which content gaps are keeping competitors in the answer set.
Caption: First-party AI search data should feed a weekly loop: observe citations, map topics, repair content, and measure again.
What changed in Bing Webmaster Tools
Bing moved first. On February 10, 2026, Microsoft introduced AI Performance in Bing Webmaster Tools as a public preview for visibility across Microsoft Copilot, AI-generated summaries in Bing, and selected partner integrations. The early dashboard focused on practical publisher questions: which pages were cited, how citation activity changed over time, and which grounding query phrases were involved.
Then on June 16, 2026, Bing expanded the preview globally with four more useful views:
| Bing signal | What it tells you | How a GEO team should use it |
|---|---|---|
| Intents | The broader purpose behind citation-triggering grounding queries, such as research, commercial, local, informational, or solve-and-learn contexts | Separate educational visibility from buyer-intent visibility instead of treating every citation as equal |
| Topics | Groups of related grounding queries into broader themes | Find topical areas where the site is becoming a trusted source, and where coverage is too thin |
| Citation Share | The percentage of citations your site receives for a grounding query out of all citations shown across sites | Watch relative presence over time without pretending it is a direct traffic-share metric |
| Compare | Time-based comparison of citation patterns | Check whether content updates, technical fixes, or market events changed AI visibility |
Caption: Bing Webmaster Tools now gives publishers a first-party AI Performance view, including citation and cited-page reporting.
The important phrase is "grounding query." It is not always the exact phrase a user typed. It is closer to the retrieval language the AI system used while looking for source material. For content teams, that difference matters. A buyer might ask, "Which SOC 2 vendor is safe for a mid-market finance team?" The grounding layer may break that into source-seeking concepts like "SOC 2 vendor comparison," "financial services compliance software," "mid-market data security requirements," and "customer support evidence."
That is where old SEO habits start to fail. A single page targeting one keyword can look good in a rank tracker and still be weak for AI retrieval because the answer needs evidence across use cases, entities, comparisons, proof points, and risk language.
What changed in Google Search Console
Google followed with a narrower but still important move. On June 3, 2026, Google announced Search Generative AI performance reports in Search Console. The reports are rolling out to a subset of websites while Google tests and gathers feedback.
The first version is not as deep as Bing's expanded AI Performance view. Google says the reports show:
| Google report field | What it means |
|---|---|
| Impressions | How often URLs from your site appeared in generative AI features in Search and Discover |
| Pages | Which URLs appeared within AI features |
| Countries | Where the visibility happened |
| Devices | Device visibility for Search results |
| Dates | Hourly, daily, weekly, and monthly performance views |
Caption: Google's Search Console report separates generative AI visibility from the broader Search performance interface.
There are obvious gaps. Most teams will want query context, click behavior, prompt patterns, answer placement, and citation quality. Google has not given all of that yet.
Still, the product decision matters. Google has separated generative AI visibility from the larger Search performance view. That is a public signal that AI search visibility is no longer a side note inside traditional SEO reporting.
The scale also explains why this matters. At Google I/O 2026, Google said AI Overviews had more than 2.5 billion monthly active users, and AI Mode had passed 1 billion monthly active users. Those numbers do not tell you how much traffic your site will get. They do tell you that AI search is large enough to deserve its own measurement model.
The real shift: from keyword ranking to retrieval coverage
The source article this piece is based on got one thing right: the industry is moving away from single-point keyword fights. The better frame is retrieval coverage.
Classic SEO asks: "Can this page rank for this query?"
GEO asks a messier question: "When an AI system decomposes a user's problem into smaller evidence needs, does our site have enough clear, trustworthy, current material to be selected as part of the answer?"
That changes the work.
A software company selling compliance automation, for example, cannot rely on a homepage plus a few feature pages. AI answers may need to understand:
- what the company does in plain language
- which customer segments it fits
- how it compares with adjacent tools
- what integrations, certifications, and support models exist
- which claims are backed by documentation or customer evidence
- whether pricing, security, implementation, and limitations are explained clearly
If those facts are scattered, stale, hidden in PDFs, or written in vague marketing language, the site may be crawlable but not very useful as an AI source.
This is why query fan-out, grounding queries, and topic clusters all point in the same direction. AI search systems often work across clusters of concepts rather than a single exact-match keyword. The content system has to match that shape.
A practical GEO measurement loop
Do not rebuild your entire content program because two dashboards changed. Start with a measured loop.
- Build an AI visibility baseline.
Export what you can from Bing Webmaster Tools and Google Search Console. Track cited pages, impressions, grounding query phrases, topics, citation share, countries, devices, and date ranges. If your Google generative AI report is not available yet, keep a placeholder in the dashboard and use Bing plus manual AI answer checks for the first pass.
- Group signals by decision journey, not by page type.
A page-level report is useful, but buyers do not think in URLs. Group signals into decision areas: definition, comparison, pricing, implementation, security, alternatives, troubleshooting, local availability, category education, and proof.
- Find the missing evidence.
A weak topic cluster usually has one of four problems: no page exists, the page exists but does not answer the sub-question, the page answers it without evidence, or the page is blocked by technical/access issues. Each problem needs a different fix.
- Repair content in small batches.
Update five to ten pages at a time. Add direct answer blocks, comparison tables, examples, source-backed claims, schema where relevant, and internal links to supporting pages. Avoid the temptation to generate dozens of thin articles. GEO rewards coherent evidence systems more than sheer publishing volume.
- Measure again after the next crawl and reporting window.
Look for movement in cited pages, topic coverage, citation share, and generative AI impressions. Treat the numbers as directional. First-party AI search reporting is young, and both Google and Bing describe parts of the experience as early or preview-stage.
Caption: A useful GEO audit maps buyer questions against current pages, evidence depth, and first-party AI visibility signals.
What most teams will get wrong
The first mistake is treating citation count as the new ranking position. It is not. Bing is careful to say Citation Share is observational, not a ranking system or competitor scoreboard. A source can be cited often because it is broadly useful, because the topic has low competition, or because the answer format happens to need that kind of source.
The second mistake is separating SEO and GEO too aggressively. You still need crawlable pages, clean internal links, useful titles, structured content, and fresh information. The difference is that those fundamentals now feed both classic search and AI answer retrieval.
The third mistake is optimizing only the pages that already appear. Existing citations show where the site has traction, but the bigger opportunity may be in topics where the brand is absent. If your product is never cited for comparison, implementation, security, or category-risk questions, that is a content strategy problem, not just a dashboard problem.
The fourth mistake is reporting screenshots as proof. Screenshots are still useful for qualitative review, especially when auditing answer wording. But a screenshot without first-party trend data is weak evidence. It captures one answer at one time, not a durable visibility pattern.
The Auspia view
The best GEO programs in 2026 will look less like content campaigns and more like data operations.
They will have a prompt library, a first-party dashboard, a topic coverage map, a page repair queue, and a review cadence. They will connect Bing's citation-level signals with Google's generative AI impressions, then use manual answer checks to understand the quality of the appearance.
Auspia's working rule is this: do not ask "Did AI mention us?" Ask "Which part of the decision journey did AI trust us for, and which part did it ignore?"
That question leads to better work. It pushes teams to clarify brand facts, publish useful supporting pages, remove vague claims, cite evidence, and build topic clusters that match how buyers actually ask questions.
If you want a quick starting point, run a visibility check with the AI Search Visibility Checker , then compare the results with your Bing Webmaster Tools and Search Console data once available. The tool will not replace first-party reporting, but it helps you find the prompts and brand-answer gaps worth monitoring.
A 7-day action plan
| Day | Action | Output |
|---|---|---|
| 1 | Export available AI search data from Bing and Google | Baseline spreadsheet |
| 2 | List the pages already appearing in AI features | Cited-page inventory |
| 3 | Group grounding queries and visible pages into topics | Topic map |
| 4 | Compare topics against the buyer journey | Gap list |
| 5 | Pick five pages to repair | Repair queue |
| 6 | Add answer blocks, evidence, tables, and supporting links | Updated pages |
| 7 | Set a weekly review cadence | GEO measurement loop |
Keep the first cycle boring. The goal is not to prove a massive win in a week. The goal is to stop guessing.
FAQ
Did Google and Bing fully open the AI search black box?
No. They opened important first-party reporting layers, but they did not expose every ranking, retrieval, click, or answer-generation factor. Bing currently gives deeper citation context, while Google's first generative AI reports focus on impressions, pages, geography, devices, and time.
Is GEO now measurable?
Yes, but not perfectly. Teams can now combine first-party AI visibility data, citation signals, page-level reports, prompt checks, and content audits. That is much better than relying on one-off screenshots, but the metrics still need careful interpretation.
Should teams stop tracking classic SEO rankings?
No. Classic SEO still matters because crawlability, authority, content quality, and user demand feed AI search visibility. The change is that rank tracking is no longer enough. You also need to know whether AI systems retrieve, cite, and summarize your content.
What is the most useful first metric to track?
Start with cited pages and generative AI impressions. Then add topic coverage, grounding query context, and citation share where available. A simple weekly dashboard beats an overbuilt model nobody uses.
What should a team fix first?
Fix pages that are already close to being useful: cited pages with thin evidence, important product pages that do not answer buyer questions, and topic gaps where competitors appear but your brand does not. Those repairs usually teach more than publishing brand-new pages from scratch.
Sources
- Bing: Introducing AI Performance in Bing Webmaster Tools Public Preview
- Bing: New AI Visibility Insights in Bing Webmaster Tools
- Google Search Central: Introducing Search Generative AI performance reports in Search Console
- Google: I/O 2026 Search updates
- Google: I/O 2026 keynote summary from Sundar Pichai
Author: Ethan Marlowe, GEO Measurement Lead Across 500+ Prompts at Auspia. Ethan writes about prompt tracking, citation reports, visibility dashboards, and practical AI search measurement for growth teams.