One page, two searches, two verdicts
Same page. Same day. Two different keywords. Two different verdicts.
SXO scan: https://auspia.ai/tools/ai_overview
SERP source: dataforseo json | organic results: 9
SERP consensus: Tool (88% - STRONG consensus)
VERDICT: ALIGNEDSXO scan: https://auspia.ai/tools/ai_overview
SERP source: dataforseo json | organic results: 8
SERP consensus: Blog (75% - STRONG consensus)
VERDICT: MISMATCH (HIGH)
recommendation: add an educational content layer (guides, docs, FAQ depth)The page did not change between the two runs. The SERP did.
The first keyword was the commercial one: "ai overview checker". Google serves nine results, eight of them tools. Our tool page fits right in. The second keyword is the one Google itself surfaced in the People Also Ask box on that first SERP: "How to check AI Overview?". Serve that phrase and Google brings back Google documentation, Reddit threads, and blog posts. Six articles, two tools. Our tool page is the wrong shape for it.
This is what search experience optimization, SXO, actually is: a check that the page type you wrote matches the page type the keyword rewards. Not the title, not the meta, not the schema. The page type. Everything else you can score 95, 96, 98 on, and the page still loses to a worse page that is the right type.
That check is now a Codex skill. This article contains the complete SKILL.md and the one scan script, tested, with the teardown above reproduced in full further down.
The symptom you keep driving past
You have seen this in some form:
What you keep seeing | What is usually actually happening |
|---|---|
The page is technically clean | Likely true. Technical health was never the problem. |
Your strongest content loses to a thinner competitor | The competitor's page is the type this keyword rewards. "Thinner" is your judgment; the SERP disagrees. |
Rankings sit at 8-13 and rewrites do nothing | Rewrites deepen a page type the keyword does not ask for. |
Clicks, then no conversions | The traffic you won was informational. You sent it to a tool page. |
The first three items in that table are the ones that break people. When a page at position 9 keeps refusing to move, the usual reflex is to rewrite it one more time. That works when the page is the right type for the keyword and the depth, schema, or authority is under par. It does not work when the SERP is saying "for this keyword I want an article" and your page is a checker.
Audits miss this because they score what is on the page. SXO scores the match between the page and the SERP. The keyword selects the page type. The page does not.
The short answer
- Who this is for: anyone with a page that ranks poorly despite good
technical health, and anyone planning content who wants to know the page type before they start writing.
- What you get: the page type of your page and of the SERP, a verdict
(ALIGNED or MISMATCH with severity), the seven-dimension gap score, and the SERP signals (People Also Ask, related searches, AI Overview) that persona cards are derived from.
- Prerequisites: python3. One SERP file: either a JSON response from a
serp/google/organic API call, or a hand-made TSV of organic results. This article includes a 6-row TSV you can run immediately, no API key.
- Definition of done: you can state, for one keyword, which type Google
rewards, your page's type, whether they match, and the largest gap. About ten minutes including the SERP capture.
Why this is a separate axis
An audit asks "is the page healthy?". SXO asks "does this page deserve to rank for this keyword?". These are different questions. A page can answer the first one with 95/100 and fail the second one completely.
Here is the thought experiment. Your SERP is 8 product pages, 2 comparison pages, zero blog posts. Your best writer spends three weeks on a definitive guide. It will not break through. Not because it is long or short, but because Google has settled what this keyword wants: products. The writer's guide is an answer to a question nobody typed.
The good news: you do not need to guess what Google settled on. The SERP itself is the data. Read it backwards, count the page types, and the reward structure is sitting right there. That is mechanical enough to automate, and the automation is exactly the boundary that makes this skill safe: the script classifies and counts, the judgment (personas, stories, wireframes) stays with the model, and every judgment must cite a signal the script extracted.
How the skill works
One gate, one taxonomy, seven dimensions, two jobs.

The gate. No verdict without data. If the SERP file has fewer than 5 organic results, the script prints a limited-confidence note and refuses to invent a consensus. If there are no organic results at all, it prints "the scan refuses to score a consensus it cannot see" and stops. No consensus, no verdict.
The taxonomy. Every page gets exactly one of eight types, using a priority order so overlaps never run silently. SERP results are classified from titles and URLs only, labeled heuristic in the report: a first pass to find the consensus, not a fact about each competitor's page.
Type | Primary signals | Typical content structure |
|---|---|---|
Landing | hero + single value prop, sign up / get started / book demo, pricing link, testimonials, minimal nav | Hero > Social proof > Features > How it works > Pricing > CTA repeat > FAQ |
Blog | byline, date, 800+ words, /blog/ path, related posts | Title > Author + Date > Intro > H2s > Examples > Conclusion > bio |
Product | price, add-to-cart / buy, multiple images, specs, reviews, SKU | Title > Images > Price + CTA > Specs > Reviews > Related |
Hybrid | education and CTAs side by side, both /blog/ and /product/ links (typical SaaS) | Problem > Solution > How it works > Deep dives > Proof > CTA > FAQ |
Service | process / methodology, case studies, team credentials, contact form | Overview > Benefits > Process > Case studies > Team > Pricing > CTA |
Comparison | "vs" or "versus", feature matrix, pros/cons, "best [X] for", review-site framing | Intro > Criteria > Matrix > Individual reviews > Verdict > FAQ |
Local | address, map embed, NAP, service area, directions | Name + location > Services > Map > Hours > Reviews > Contact |
Tool | input fields, calculator / generator / checker, utility-first, "free [name]" | Tool above fold > Brief instructions > Output > Supporting content > FAQ |
Classification priority: interactive tool actually present → Tool; physical address + map → Local; comparison table + "vs" framing → Comparison; price + buy button → Product; CTA-heavy + minimal navigation → Landing; service process + case studies → Service; education + CTA mix → Hybrid; default → Blog.
The consensus thresholds. One number decides the verdict, and the verdict decides the recommendation:
Dominant page type share | Verdict |
|---|---|
> 60% | STRONG consensus: this is the page type Google rewards |
40-60% | MIXED: two clusters of intent; pick the cluster with more volume |
< 40% | FRAGMENTED: no dominant type - differentiation opportunity |
The severity table. A mismatch is not uniform. Being the wrong type in a product SERP costs more than being almost right in a mixed one:
Target type | SERP expects | Severity | Recommended move |
|---|---|---|---|
Blog Post | Product / Tool / Local / Landing | CRITICAL | Build the other page type; the post will not break through |
Blog Post | Comparison | HIGH | Restructure as a comparison with a matrix |
Product | Informational (Blog) | HIGH | Add an educational content layer |
Tool | Informational (Blog) | HIGH | Add an educational content layer (guides, docs, FAQ depth) |
Landing | Tool | HIGH | Build the interactive component |
Service | Local results | MEDIUM | Add location signals + LocalBusiness schema |
any | matching | ALIGNED | Focus on content depth and UX |
The seven dimensions. Score the page against the SERP's expectations, not against a checklist. Freshness is 10 points, the other six are 15:
Dimension | Compare | Points |
|---|---|---|
Page type | target vs SERP dominant type | 0-15 |
Content depth | word count, H2 coverage vs expected depth | 0-15 |
UX signals | CTA clarity, above-fold answer, mobile | 0-15 |
Schema markup | present vs expected structured data | 0-15 |
Media richness | images, video, interactivity vs SERP norm | 0-15 |
Authority signals | author, proof, credentials, reviews | 0-15 |
Freshness | last updated, date signals | 0-10 |
Two jobs. Mechanics are the script's job. Judgment is yours: the script extracts the signals, and persona cards, user stories, and wireframes must cite those signals. That division is the whole safety property of the skill. The script never invents a persona; it hands the model the People Also Ask questions, related searches, ad themes, and AI Overview citations, and says "these are your sources."
The full SKILL.md
Everything here is the verbatim skill file. Tasked with "my page is optimized but not ranking", this is what runs:
---
name: codex-seo-sxo
description: Use when the user asks about search experience optimization, SXO, page type mismatch, intent mismatch, why a page is not ranking despite being well optimized, SERP analysis against one page, user stories from search intent, persona-based page scoring, SERP consensus, or IST/SOLL wireframes. Classifies a page and a SERP into the eight page types, detects the mismatch, and scores the seven-dimension gap.
---
# Search Experience Optimization (SXO)
SXO bridges SEO (what Google rewards) and UX (what users need). The core
question is not "is this page technically healthy" but "does this page
deserve to rank for this keyword based on what the SERP actually rewards".
## Core insight
A page can score 95/100 on technical SEO and still fail to rank because it
is the **wrong page type** for the keyword. If Google shows 8 product pages
and 2 comparison pages for your keyword, your blog post will never break
through - no matter how well written. The keyword selects the page type;
the page does not.
## Two jobs
1. Classify the target page and the SERP into the eight page types,
detect a mismatch, and score the seven-dimension SXO gap.
2. Derive user stories and persona scores from the SERP's own signals
(PAA questions, related searches, ad themes, AI Overview) so every fix
names a person, a goal, and a barrier.
Mechanics are the script's job. Judgment is yours: the script extracts the
signals; persona cards and user stories must cite those signals.
## Commands
```bash
python3 ~/.codex/skills/codex-seo-sxo/scripts/sx_scan.py <page-url> <serp.json|serp.tsv>
python3 ~/.codex/skills/codex-seo-sxo/scripts/sx_scan.py <page-url> <serp.json> --json
```
The SERP file is a capture, never a scrape. Copy the response of a
`serp/google/organic` style API call (`serp_organic_live` family, saved as
JSON) or build a simple TSV with one row per organic result:
```
title<TAB>url
```
The target page is fetched directly (a local HTML file also works when bot
checks block the script). Nothing here talks to Google. If you cannot get
a SERP file with at least 5 organic results, say so and do not invent a
consensus.
## Page-type taxonomy
Every page - target and each SERP result - gets exactly one of eight
types. When signals overlap, use the priority order at the end.
| Type | Primary signals | Typical content structure |
|------|-----------------|---------------------------|
| Landing | hero + single value prop, sign up / get started / book demo, pricing link, testimonials, minimal nav | Hero > Social proof > Features > How it works > Pricing > CTA repeat > FAQ |
| Blog | byline, date, 800+ words, /blog/ path, related posts | Title > Author + Date > Intro > H2s > Examples > Conclusion > bio |
| Product | price, add-to-cart / buy, multiple images, specs, reviews, SKU | Title > Images > Price + CTA > Specs > Reviews > Related |
| Hybrid | education and CTAs side by side, both /blog/ and /product/ links (typical SaaS) | Problem > Solution > How it works > Deep dives > Proof > CTA > FAQ |
| Service | process / methodology, case studies, team credentials, contact form | Overview > Benefits > Process > Case studies > Team > Pricing > CTA |
| Comparison | "vs" or "versus", feature matrix, pros/cons, "best [X] for", review-site framing | Intro > Criteria > Matrix > Individual reviews > Verdict > FAQ |
| Local | address, map embed, NAP, service area, directions | Name + location > Services > Map > Hours > Reviews > Contact |
| Tool | input fields, calculator / generator / checker, utility-first, "free [name]" | Tool above fold > Brief instructions > Output > Supporting content > FAQ |
**Classification priority** (use the strongest signal first):
1. interactive tool actually present: **Tool**
2. physical address + map: **Local**
3. comparison table + "vs" framing: **Comparison**
4. price + buy button: **Product**
5. CTA-heavy + minimal navigation: **Landing**
6. service process + case studies: **Service**
7. education + CTA mix: **Hybrid**
8. default: **Blog**
SERP results are classified from titles and URLs only - label that
heuristic basis in the report. It is a first pass to find the consensus,
not a fact about each competitor's page.
## SERP consensus
| Dominant page type share | Verdict |
|--------------------------|---------|
| > 60% | STRONG consensus: this is the page type Google rewards |
| 40-60% | MIXED: two clusters of intent; pick the cluster with more volume |
| < 40% | FRAGMENTED: no dominant type - differentiation opportunity |
Never report a consensus from fewer than 5 organic results; note limited
confidence instead.
## Mismatch severity
| Target type | SERP expects | Severity | Recommended move |
|-------------|--------------|----------|------------------|
| Blog Post | Product / Tool / Local / Landing | CRITICAL | Build the other page type; the post will not break through |
| Blog Post | Comparison | HIGH | Restructure as a comparison with a matrix |
| Product | Informational (Blog) | HIGH | Add an educational content layer |
| Tool | Informational (Blog) | HIGH | Add an educational content layer (guides, docs, FAQ depth) |
| Landing | Tool | HIGH | Build the interactive component |
| Service | Local results | MEDIUM | Add location signals + LocalBusiness schema |
| any | matching | ALIGNED | Focus on content depth and UX |
A fragmented SERP is not a mismatch - note the differentiation
opportunity instead. The SXO verdict is separate from any technical
health score; report both when both exist and say which is which.
## Seven-dimension gap score
Score the page against the SERP's expectations, not against a checklist:
| Dimension | Compare | Points |
|-----------|---------|--------|
| Page type | target vs SERP dominant type | 0-15 |
| Content depth | word count, H2 coverage vs expected depth | 0-15 |
| UX signals | CTA clarity, above-fold answer, mobile | 0-15 |
| Schema markup | present vs expected structured data | 0-15 |
| Media richness | images, video, interactivity vs SERP norm | 0-15 |
| Authority signals | author, proof, credentials, reviews | 0-15 |
| Freshness | last updated, date signals | 0-10 |
Lower = larger gap. The script prints its proxy for every point with the
evidence it saw; when a proxy is not the real measure (for example no
date on an evergreen tool page), the note says so and you re-judge by
hand.
## User stories from signals
Never guess. Cite the signal that generated each story:
| Signal | What it reveals | Example persona |
|--------|-----------------|-----------------|
| PAA question cluster | knowledge gaps and concerns | Beginner, Technical Evaluator, Skeptic |
| Ad copy themes | commercial triggers and objections | Budget Buyer, Risk-Averse Decision Maker |
| Related searches | journey before/after | Researcher, Comparison Shopper |
| Featured snippet format | expected answer structure | Quick-Answer Seeker (paragraph), Steps Person (list) |
| AI Overview citations | the synthesis Google already trusts | AI-Informed Reader |
Write 3-5 stories in the format: *As a [persona], I want to [goal],
because [emotional driver], but I'm blocked by [barrier].* The barrier
should be specific enough to suggest a page-level fix (information gap,
trust gap, comparison fatigue, price sensitivity, technical confusion,
time pressure).
## Persona scoring (4-7 personas)
Score the page from each persona's perspective, 25 points per dimension:
| Range | Relevance (0-25) | Clarity (0-25) | Trust (0-25) | Action (0-25) |
|-------|------------------|----------------|--------------|---------------|
| 21-25 | directly addresses the persona's primary goal | answer above fold for this persona | multiple persona-relevant trust signals | persona-appropriate CTA, low friction |
| 16-20 | covers topic, lacks persona depth | answer on page, needs scroll | general trust signals | generic CTA |
| 11-15 | tangential - persona must extrapolate | buried in dense text | some signals, real gaps | next step exists but hidden |
| 6-10 | serves a different audience | pieced together from sections | minimal signals | one CTA that does not match the stage |
| 0-5 | irrelevant | no clear answer | none, or undermining signals | dead end |
Every persona must trace to a SERP signal - no invented personas. Weight
personas by intent share (ads dominate = commercial personas up; AI
Overview = the AI-informed reader). Rank fixes: weakest persona with the
highest volume weight first, then the weakest dimension across personas
(systemic issue), then critical mismatches.
## Wireframe (optional)
For a rebuild, produce an IST (current state) and a SOLL (target state)
semantic HTML outline, mobile first: the largest fold (about 600px at
375px wide) must carry the single most important element for the page
type. Placeholders must be concrete - never "add a CTA" but "add pricing
CTA with annual savings badge below the hero, linking to /pricing#enterprise".
The SOLL mirrors the SERP-expected page type, ordered per its content
structure in the taxonomy table.
## Output contract
Report the seven sections: SERP landscape (dominant type + confidence +
features), page-type alignment (target vs expected + verdict), user
stories with cited signals, gap table with the total, persona cards with
the four scores, priority actions (mismatch first, then weakest persona),
and limitations (what could not be assessed and why). State the data
source and its age. If the SERP capture had no AI Overview text, say
"not returned (async render)" rather than claiming its absence.
## Errors
| Scenario | Action |
|----------|--------|
| Page fetch fails (DNS, timeout, 403) | Report the failure; never score a cached copy |
| No SERP file, or < 5 organic results | Refuse the consensus; ask for a capture or TSV |
| `--json` on empty SERP | Output the honesty gate message, no score |
| All paid results | Note the commercial SERP; analyze ad copy only |
| JS-rendered target page | Fetch the rendered HTML via the user and rerun |
| Mismatch is fragmented | Note opportunity, skip severity |
## Quality checklist
- [ ] At least 5 SERP results analyzed; count given
- [ ] Page type from page signals; SERP classification labeled heuristic
- [ ] User stories cite specific SERP signals as evidence
- [ ] Personas all trace to evidence; 4-7, not more
- [ ] Weakest persona treated first in recommendations
- [ ] SXO score clearly separate from technical health score
- [ ] Limitations section present and honest
## Security
No credentials stored or transmitted. The scripts fetch only the target
URL or file provided and read the SERP file you give them; they never
scrape Google or call any API by themselves.The scan script
That 484-line file sits in scripts/sx_scan.py. Standard library only: a fetch with gzip and a user agent, a body-to-text extractor, the classification functions, a DataForSEO JSON / TSV reader, the mechanical gap scorer, and the printer. Here is the whole thing:
#!/usr/bin/env python3
"""Search Experience Optimization scan: classify a page and a SERP into the
eight page types, detect a page-type mismatch, and score the seven-dimension
SXO gap against the SERP's expectations.
The SERP is never scraped here: it must be provided as a file (a
DataForSEO serp/google/organic/live/advanced response saved as JSON, or a
simple TSV of organic results with one row per URL). The script scores only
what it can see - classification of SERP results is title/URL heuristics,
and the persona cards, user stories and wireframe are the model's job
after this scan returns the material.
Usage: python3 sx_scan.py <page-url-or-file> <serp.json|serp.tsv> [--json]
"""
import json
import os
import re
import sys
import urllib.error
import urllib.request
UA = ("Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
"AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Safari/537.36")
TOOL_WORDS = ("checker", "calculator", "generator", "analyzer", "tracker",
"tool", "simulator", "scanner", "validator", "converter")
COMPARE_WORDS = ("vs", "versus", "comparison", "compare", "best ",
"top ", "list", "review", "alternatives")
PRODUCT_WORDS = ("buy", "price", "shop", "free shipping", "add to cart",
"in stock")
LOCAL_WORDS = ("near me", "address", "map", "directions", "hours of",
"opening hours", "location")
BLOG_DOMAINS = ("support.google.com", "developers.google.com", "wikipedia.org",
"reddit.com", "blog.google", "hubspot.com", "neilpatel.com",
"semrush.com/blog", "moz.com/blog", "search."
"enginejournal.com", "searchengineland.com", "quora.com",
"medium.com", "capterra.com/blog")
TOOL_DOMAINS = ("smallseotools.com", "duplichecker.com", "sitechecker.pro",
"indexcheckr", "indexly.ai", "indexchecker.io",
"zenserp.com", "semrush.com/free-tools", "seo.com/tools",
"sitechecker", "www.seo.com", "seomator.com",
"pikaseo.com", "growthnatives.com",
"aeoengine.ai", "hnusolutions.com", "advancedwebranking.com")
COMPARE_DOMAINS = ("g2.com", "getapp.com", "trustradius.com",
"capterra.com", "softwareadvice.com")
def fetch(source):
if os.path.exists(source):
return open(source, encoding="utf-8", errors="replace").read(), "file"
url = source if "://" in source else "https://" + source
try:
req = urllib.request.Request(url, headers={"User-Agent": UA,
"Accept-Encoding": "gzip"})
with urllib.request.urlopen(req, timeout=25) as r:
raw = r.read()
except urllib.error.HTTPError as e:
raise SystemExit("FETCH FAIL: %s HTTP %s" % (url, e.code))
except Exception as e:
raise SystemExit("FETCH FAIL: %s (%s)" % (url, e))
if raw[:2] == b"\x1f\x8b":
import gzip
raw = gzip.decompress(raw)
return raw.decode("utf-8", "replace"), url
def visible(html):
"""Plain text of the page body, scripts and styles removed."""
m = re.search(r"(?s)<body\b.*?</body>", html)
body = m.group(0) if m else html
body = re.sub(r"(?s)<(script|style|noscript)\b.*?</\1>", " ", body)
main = re.search(r"(?s)<main\b.*?</main>", body)
if main:
body = main.group(0)
text = re.sub(r"<[^>]+>", " ", body)
return re.sub(r"\s+", " ", text).strip()
def find_attrs(html, path):
"""Count occurrences of an attribute pattern in img/input-like tags."""
return len(re.findall(path, html, re.I))
# ---------------------------------------------------------------- taxonomy
def classify_page(html, text, title):
"""Lean on the classification priority from page-type-taxonomy.md."""
h1 = re.search(r"(?s)<h1\b[^>]*>(.*?)</h1>", html)
h1 = re.sub(r"<[^>]+>", " ", h1.group(1)).strip() if h1 else ""
head = (title + " " + h1 + " " + text[:800]).lower()
interactive = bool(re.search(r"<input\b", html, re.I)) and (
bool(re.search(r"<form\b", html, re.I)) or
bool(re.search(r'role="search"', html, re.I)) or
bool(re.search(r"<button\b", html, re.I)))
schema = [s for s in re.findall(r'"@type"\s*:\s*"([^"]+)"', html)]
has_webapp = "WebApplication" in schema or "SoftwareApplication" in schema
is_tool = ((has_webapp or interactive) and
any(w in head for w in TOOL_WORDS))
if is_tool:
return "Tool"
has_address = bool(re.search(r"\b[0-9]{1,5}\s+[A-Z][A-Za-z']*(\s|$).*?"
r"\b(ave|st|street|blvd|rd|way|ln|dr)\b",
text, re.I)) or \
bool(re.search(r"LocalBusiness|PostalAddress", html)) or \
("maps.google" in html or "google.com/maps" in html) and \
any(w in text.lower() for w in ("directions", "our store", "visit us"))
if has_address or "LocalBusiness" in schema:
return "Local"
has_table = bool(re.search(r"<table\b", html, re.I)) and \
re.search(r"<th\b|<thead\b", html, re.I)
if has_table and any(w in head for w in COMPARE_WORDS) or \
(re.search(r"\bvs\.?\b", head, re.I) and has_table):
return "Comparison"
if "$" in text and any(w in head for w in PRODUCT_WORDS) or \
("Product" in schema and "price" in text.lower()):
return "Product"
cta_words = ("get started", "sign up", "book a demo", "book a call",
"start free", "subscribe", "try it free", "pricing",
"buy now", "contact us", "free trial")
ctas = sum(1 for w in cta_words if w in head)
nav = re.findall(r"<nav\b[^>]*>(.*?)</nav>", html, re.I | re.S)
nav_links = len(re.findall(r"<a\b", " ".join(nav), re.I)) if nav else 99
if ctas >= 2 and nav_links <= 10:
return "Landing"
process = bool(re.search(r"how we work|our process|methodology|case "
r"studies|portfolio", text.lower()))
contact = bool(re.search(r"<form\b", html, re.I)) or \
bool(re.search(r"contact\b|book a consult", text.lower()))
if process and contact:
return "Service"
edu_ctas = ctas >= 1 and bool(re.search(r"<article\b|/blog|author|"
r"related posts", html, re.I))
if edu_ctas or (has_webapp is False and has_table is False and
len(schema) == 0):
return "Hybrid" if ctas >= 1 and not has_webapp else "Blog"
return "Blog"
def classify_serp(title, url):
"""Heuristic classification of a SERP result from its title and URL.
Label it heuristic in the report - it is a first pass, not a verdict."""
t = (title or "").lower()
u = (url or "").lower()
if any(dom in u for dom in BLOG_DOMAINS):
return "Blog"
if any(dom in u for dom in COMPARE_DOMAINS):
return "Comparison"
# Listicle/comparison patterns win over tool words: "Best X Tools",
# "Top X", "X Comparison", "vs". Otherwise "AI Overview Checker" style
# titles misclassify as Comparison and "Best ... Checkers" as Tool.
if "versus" in t or re.search(r"\b\w+\s+vs\s+\w+\b", t) or \
any(w in t for w in ("comparison", "compare")) or \
(("best " in t or "top " in t or "list" in t) and
any(w in t for w in ("tools", "tool", "software", "alternatives",
"checkers", "checker", "plugins"))):
return "Comparison"
if any(dom in u for dom in TOOL_DOMAINS) or \
("/tools/" in u or "/free-tools/" in u or "/tool/" in u) or \
any(w in t for w in TOOL_WORDS):
return "Tool"
if any(w.lower() in t for w in ("price", "buy", "shop", "order now")) or \
"/product" in u:
return "Product"
if any(w in t.lower() for w in ("near me", "in ", "at ", "location")) and \
re.search(r"\b(city|town|venue|plaza|center)\b", t):
return "Local"
return "Blog"
# ------------------------------------------------------------------- input
def load_serp(path):
"""Accept a DataForSEO response JSON or a simple TSV. Returns
(organics, paa, related, ads, aio_md) where organics is a list of
dicts {title, url} in SERP rank order."""
if path.endswith(".tsv") or "\t" in open(path, encoding="utf-8",
errors="replace").read()[:200]:
org = []
for line in open(path, encoding="utf-8", errors="replace"):
parts = line.rstrip("\n").split("\t")
if len(parts) >= 2:
org.append({"title": parts[0].strip() or "",
"url": parts[1].strip()})
return {"organic": org, "paa": [], "related": [], "ads": [],
"aio_md": "", "aio_seen": False, "src": "tsv"}
data = json.load(open(path, encoding="utf-8"))
tasks = data.get("tasks", [])
if not tasks or tasks[0].get("status_code") not in (20000, 20101):
raise SystemExit("SERP FILE: tasks[0].status_code=%s %s" % (
tasks[0].get("status_code"), tasks[0].get("status_message")))
r = tasks[0]["result"][0]
org, paa, related, ads, aio_md, aio_seen = [], [], [], [], "", False
for it in r.get("items", []):
ty = it.get("type", "")
if ty == "organic":
org.append({"title": (it.get("title") or "").strip(),
"url": (it.get("url") or "").strip()})
elif ty == "people_also_ask":
paa += [e.get("title") for e in it.get("items", [])
if e.get("title")]
elif ty == "related_searches":
related += [s for s in it.get("items", []) if isinstance(s, str)]
elif ty == "paid":
ads.append((it.get("title") or ""))
elif ty == "ai_overview":
aio_seen = True
aio_md = it.get("markdown") or ""
return {"organic": org, "paa": paa, "related": related, "ads": ads,
"aio_md": aio_md, "aio_seen": aio_seen, "src": "dataforseo json"}
# --------------------------------------------------------------------- gap
def score_gap(html, text, page_type, expected):
"""Seven dimensions, 0-15 each except Freshness (0-10). Mechanical
proxy scoring: every point is explained in the margins so a human can
re-judge the ones where the proxy is wrong."""
rows = []
schema = [s for s in re.findall(r'"@type"\s*:\s*"([^"]+)"', html)]
h2 = len(re.findall(r"<h2\b", html, re.I)) + 1
words = len(text.split())
# 1. Page type
if page_type == expected:
pt, note = 15, "page type matches SERP consensus"
elif expected == "fragmented":
pt, note = 12, "no dominant type - differentiation opportunity"
else:
sev = {"CRITICAL": 0, "HIGH": 5, "MEDIUM": 10}.get(
mismatch_severity(page_type, expected), 5)
pt, note = sev, ("%s mismatch severity %s" % (
page_type, mismatch_severity(page_type, expected)))
rows.append(("Page type", "15", pt, note))
# 2. Content depth (proxy: words + H2 density)
dw = 15 if words >= 1200 else 10 if words >= 600 else 6 if words >= 300 else 2
dh = 3 if h2 >= 4 else 1 if h2 >= 2 else 0
rows.append(("Content depth", "15", min(15, dw + dh),
"%d words, %d H2 headings" % (words, h2)))
# 3. UX signals (proxy: viewport, H1, visible action above fold)
ux = 0
uxn = []
if re.search(r'<meta\b[^>]*viewport', html, re.I):
ux += 6
uxn.append("viewport meta")
if re.search(r"<h1\b", html, re.I):
ux += 4
uxn.append("one h1")
if re.search(r"<(input|textarea|button)\b", html[:6000], re.I) or \
re.search(r'placeholder="[^"]{5,}', html[:3000], re.I):
ux += 5
uxn.append("action element in first screen")
rows.append(("UX signals", "15", min(15, ux),
", ".join(uxn) if uxn else "no measurable UX signals"))
# 4. Schema markup
expected_schema = {"Tool": "WebApplication", "Blog": "Article",
"Product": "Product", "Comparison": "ItemList",
"Local": "LocalBusiness", "Service": "Service",
"Landing": "SoftwareApplication",
"Hybrid": "SoftwareApplication"}.get(page_type, "")
sx = 0
sxn = []
if expected_schema and expected_schema in schema:
sx += 8
sxn.append("%s present" % expected_schema)
elif expected_schema:
sxn.append("%s missing" % expected_schema)
if len(schema) and re.search(r'"@context"\s*:\s*"https://schema\.org"',
html):
sx += 4
sxn.append("valid JSON-LD @context")
if re.search(r'FAQPage', html) and page_type in ("Tool", "Landing"):
sx += 3
sxn.append("FAQPage block")
rows.append(("Schema markup", "15", min(15, sx),
", ".join(sxn) if sxn else "no schemas detected"))
# 5. Media richness
imgs = len(re.findall(r"<img\b", html, re.I))
alts = len(re.findall(r'<img\b[^>]*alt="[^"]{3,}"', html, re.I))
mx = 0
mxn = []
if imgs >= 3:
mx += 5
mxn.append("%d images" % imgs)
if imgs and alts >= imgs * 0.5:
mx += 5
mxn.append("alt coverage %d/%d" % (alts, imgs))
if re.search(r"<video\b", html, re.I):
mx += 5
mxn.append("video present")
rows.append(("Media richness", "15", max(0, min(15, mx)),
", ".join(mxn) if mxn else ("%d images, no video" % imgs)))
# 6. Authority signals (proxy: org schema, author, testimonials, refs)
ax = 0
axn = []
if "Organization" in schema or "WebSite" in schema:
ax += 5
axn.append("Organization/WebSite schema")
if re.search(r"author|byline|written by|owner", text, re.I) or \
"Author" in html:
ax += 4
axn.append("author entity")
if re.search(r"testimonial|\breviews?\b|trusted by|customer story|logo",
text, re.I):
ax += 3
axn.append("social proof phrases")
if re.search(r"2[0-9]{3}|20[0-9]{2}", html):
ax += 3
axn.append("years / numbers")
rows.append(("Authority signals", "15", min(15, ax),
", ".join(axn) if axn else "no authority signals found"))
# 7. Freshness (proxy: dates anywhere on page)
fx = 0
fxn = []
if re.search(r'datePublished|dateModified', html):
fx += 6
fxn.append("datePublished/dateModified")
if re.search(r"updated|last published|\bupdated\b", text, re.I) and fx == 0:
fx += 4
fxn.append("visible 'updated' text")
rows.append(("Freshness", "10", min(10, fx),
", ".join(fxn) if fxn else "no date signals (tool page " +
"may legitimately not need one - judgment call)"))
return rows
def mismatch_severity(page_type, expected):
table = {("Blog", "Product"): "CRITICAL", ("Blog", "Tool"): "CRITICAL",
("Blog", "Comparison"): "HIGH", ("Product", "Blog"): "HIGH",
("Product", "Tool"): "HIGH",
("Tool", "Blog"): "HIGH", ("Tool", "Comparison"): "HIGH",
("Landing", "Tool"): "HIGH", ("Service", "Local"): "MEDIUM",
("Landing", "Blog"): "CRITICAL", ("Blog", "Local"): "CRITICAL",
("Blog", "Landing"): "CRITICAL"}
return table.get((page_type, expected), "MEDIUM")
def consensus(classes):
if not classes:
return "", 0, "no organic results"
from collections import Counter
c = Counter(classes)
dom, n = c.most_common(1)[0]
pct = 100.0 * n / len(classes)
if pct > 60:
level = "STRONG consensus"
elif pct >= 40:
level = "MIXED consensus"
else:
level = "FRAGMENTED (no dominant type)"
return dom, pct, level
# ------------------------------------------------------------------ output
def main():
args = [a for a in sys.argv[1:] if not a.startswith("--")]
want_json = "--json" in sys.argv
if len(args) < 2:
print("sx_scan.py - SXO scan (claude-seo-derived, Codex edition)")
print(" python3 sx_scan.py <page-url|file.html> <serp.json|serp.tsv>")
print(" python3 sx_scan.py <page> <serp.json> --json")
print(" The SERP file is a DataForSEO serp/google/organic/live/")
print(" advanced response, or a TSV: title<TAB>url per row.")
return
html, source = fetch(args[0])
text = visible(html)
title = (re.search(r"<title[^>]*>(.*?)</title>", html, re.S)
or [None, ""])[1]
title = re.sub(r"<[^>]+>", " ", title).strip() if title else ""
page_type = classify_page(html, text, title)
serp = load_serp(args[1])
org = serp["organic"]
if len(org) < 5:
print(" ! sample note: only %d organic results in the SERP file; "
"consensus confidence is limited" % len(org))
serp_classes = [(classify_serp(r["title"], r["url"])) for r in org]
dom, pct, level = consensus(serp_classes)
if not org:
result = {"target": args[0], "page_type": page_type,
"consensus": "no data", "verdict": "NO SERP DATA"}
print("SXO scan: %s" % args[0])
print(" page type: %s" % page_type)
print(" VERDICT: no organic results in the SERP file - the scan "
"refuses to score a consensus it cannot see.")
return
# one keyword, one expected type
if dom:
sev = mismatch_severity(page_type, dom)
verdict = "ALIGNED" if page_type == dom else (
"MISMATCH (%s)" % sev)
rows = score_gap(html, text, page_type, dom)
header = ["#", "Domain", "Title", "Class"]
print("SXO scan: %s" % args[0])
print(" target page type: %s (from page signals)" % page_type)
print(" SERP source: %s | organic results: %d" % (
serp["src"], len(org)))
print(" SERP consensus: %s (%d%% - %s)" % (
dom, pct, level))
print(" VERDICT: %s" % verdict)
if page_type != dom and dom:
print(" recommendation: %s" % RECS.get(page_type + ">" + dom,
"reconsider page type"))
print("\n SERP result classification (heuristic - titles and URLs only):")
for i, (r, c) in enumerate(zip(org, serp_classes), 1):
print(" %2d %-22s %s" % (i, c, (r["url"] or "")[:60]))
counts = {k: serp_classes.count(k) for k in sorted(set(serp_classes))}
print(" types: %s" % " | ".join("%s %d" % kv for kv in counts.items()))
if serp["aio_md"]:
cited = sorted(set(re.findall(r"https?://([a-z0-9.-]+)", serp["aio_md"])))
print(" AI Overview: present (%d chars), citations incl. %s" % (
len(serp["aio_md"]), ", ".join(cited or ["(none)"])))
elif serp.get("aio_seen"):
print(" AI Overview: element present in the capture but no markdown "
"returned (asynchronous render) - treat as present")
else:
print(" AI Overview: not present in this capture")
if serp["paa"]:
print(" PAA questions (%d):" % len(serp["paa"]))
for q in serp["paa"]:
print(" - %s" % q)
print(" Related searches (%d): %s" % (
len(serp["related"]), " | ".join(serp["related"])[:250]))
print(" Ads on capture: %d" % len(serp["ads"]))
print("\n SXO gap score (separate from any technical-health score):")
total = 0
for name, max_, got, note in rows:
total += int(got)
print(" %-18s %2s/%2s %s" % (name, got, max_, note))
print(" TOTAL: %d/100" % total)
if want_json:
print(json.dumps({"target": args[0], "page_type": page_type,
"serp_source": serp["src"],
"organic_count": len(org),
"consensus": {"type": dom, "pct": pct,
"level": level},
"verdict": verdict,
"gap_score": {"total": total,
"rows": rows},
"paa": serp["paa"],
"related": serp["related"],
"aio_citations": (
sorted(set(re.findall(
r"https?://([a-z0-9.-]+)", serp["aio_md"])))
if serp["aio_md"] else [])},
indent=2))
RECS = {
"Blog>Product": "create a dedicated product page for this keyword",
"Blog>Tool": "build the interactive tool - a blog post will not break through",
"Blog>Comparison": "restructure as a comparison with a matrix",
"Blog>Local": "add location signals and LocalBusiness schema",
"Blog>Landing": "write a landing page, not another post",
"Product>Blog": "add an educational content layer to the product page",
"Tool>Blog": "add an educational content layer (guides, docs, FAQ depth)",
"Tool>Comparison": "add a comparison/alternatives section",
"Landing>Tool": "build the interactive tool component",
"Landing>Blog": "this keyword is informational - do not force a landing page",
"Service>Local": "add location signals plus LocalBusiness schema",
"Tool>Landing": "the SERP wants a sign-up page; keep tool behind the CTA",
}
if __name__ == "__main__":
main()Install it in two commands
Create the directories and paste the two blocks:
mkdir -p ~/.codex/skills/codex-seo-sxo/scripts- Save the
~~~~markdownblock as
~/.codex/skills/codex-seo-sxo/SKILL.md.
- Save the
~~~~pythonblock as
~/.codex/skills/codex-seo-sxo/scripts/sx_scan.py.
- Save this smoke-test file as
~/.codex/skills/codex-seo-sxo/scripts/smoke.tsv: six real organic results from the "ai overview checker" SERP shown later, one title<TAB>url per line:
Free AI Overview Checker | Track your site in AI results! https://www.seo.com/tools/ai-overview-checker/
Google AI Overview Checker & Tracker https://sitechecker.pro/google-ai-overview/
Free AI Overview Checker – Score Your GEO Readiness 2026 https://hnusolutions.com/ai-overview-checker/
Google AI Overview Keywords Checker FREE https://seomator.com/google-ai-overview-keywords-checker
AI Overviews Checker: Track Your Visibility in Google's AIOs https://www.semrush.com/free-tools/ai-overviews-visibility-checker/
Best Google AI Overviews Checker Tools in 2026 (Free ... https://llmpulse.ai/blog/best-ai-overviews-checkers/Then run:
python3 ~/.codex/skills/codex-seo-sxo/scripts/sx_scan.py https://auspia.ai/tools/ai_overview ~/.codex/skills/codex-seo-sxo/scripts/smoke.tsvYou should see SERP consensus: Tool and VERDICT: ALIGNED before the seven-dimension table and the gap total. If python3 is missing or the file was pasted wrong, you see it on the first line instead. Then run it against your own page and a SERP file of your own (the 8-9 rows for your keyword: from a SERP API's JSON response, or titles and URLs you copy out of a manual search you ran yourself).
The target page is fetched directly: a local HTML file also works when bot checks block the script. Nothing in the skill talks to Google.
The diagnosis, step by step
The page under the lens
The page is our own AI Overview citation checker at https://auspia.ai/tools/ai_overview. You type a keyword and it lists which domains Google's AI Overview cites for that keyword. From the scan: 1011 words, 12 H2 headings, one H1 ("Which domains does AI cite for your keyword?"), an input + button above the fold, WebApplication plus FAQPage plus Organization schema, 4 images with 4/4 alt coverage, no dates anywhere.
Two SERP captures were taken on 2026-08-29 (DataForSEO, serp/google/organic/live/advanced, United States, English, desktop). The keyword for run one is the page's own commercial head term. The keyword for run two is a People Also Ask question from that first SERP, which is exactly how the skill would hand you a sibling keyword: "what else does my SERP reveal, and is my page shaped for that too?"
Run one: the commercial SERP
SXO scan: https://auspia.ai/tools/ai_overview
target page type: Tool (from page signals)
SERP source: dataforseo json | organic results: 9
SERP consensus: Tool (88% - STRONG consensus)
VERDICT: ALIGNED
SERP result classification (heuristic - titles and URLs only):
1 Tool https://www.seo.com/tools/ai-overview-checker/
2 Tool https://sitechecker.pro/google-ai-overview/
3 Tool https://hnusolutions.com/ai-overview-checker/
4 Tool https://seomator.com/google-ai-overview-keywords-checker
5 Tool https://www.semrush.com/free-tools/ai-overviews-visibility-c
6 Tool https://aioverview.growthnatives.com/ai-overview-visibility-
7 Comparison https://llmpulse.ai/blog/best-ai-overviews-checkers/
8 Tool https://www.advancedwebranking.com/free-seo-tools/google-ai-
9 Tool https://pikaseo.com/free-tools/ai-overview-analyzer
types: Comparison 1 | Tool 8
AI Overview: element present in the capture but no markdown returned (asynchronous render) - treat as present
PAA questions (4):
- How to check AI Overview?
- Is there a 100% accurate AI detector?
- Can I get rid of AI on my Google Search?
- Is AI Overview ChatGPT?
Related searches (8): Ai overview checker free | Best ai overview checker | Ai overview checker online free | Ai overview checker online | Ai overview checker google | AI Overview Google | AI Overview Search | AI checker
Ads on capture: 0
SXO gap score (separate from any technical-health score):
Page type 15/15 page type matches SERP consensus
Content depth 13/15 1011 words, 12 H2 headings
UX signals 10/15 viewport meta, one h1
Schema markup 15/15 WebApplication present, valid JSON-LD @context, FAQPage block
Media richness 10/15 4 images, alt coverage 4/4
Authority signals 15/15 Organization/WebSite schema, author entity, social proof phrases, years / numbers
Freshness 0/10 no date signals (tool page may legitimately not need one - judgment call)
TOTAL: 78/100Read it the way the skill wants you to read it. The verdict is ALIGNED: 8 of 9 results are Tool pages, so Google rewards a tool here, and our page is a tool. The gap column is then graded against the same SERP, not against an abstract checklist: Page type 15/15, and the four dimensions the page actually performs well on (content depth, schema, media, authority) already carry most of their weight.
Two details to see with fresh eyes. The llmpulse result at #7 is a "Best AI Overviews Checkers" listicle, and the classifier called it Comparison deliberately: listicle patterns ("best ... checkers") win over tool words in the classification order, and the report shows the call so you can disagree with it. And the AI Overview line says "element present in the capture but no markdown returned (asynchronous render) — treat as present" rather than claiming the AI Overview does not exist. The script keeps honest, mechanically.
The two findings the script did not print: the PAA question "How to check AI Overview?" is the gateway to run two, and related searches like "Best ai overview checker" and "Ai overview checker free" run from the commercial into the comparison lane. Those are persona material, and personas are the model's job. They come later.
Run two: the PAA sibling
SXO scan: https://auspia.ai/tools/ai_overview
target page type: Tool (from page signals)
SERP source: dataforseo json | organic results: 8
SERP consensus: Blog (75% - STRONG consensus)
VERDICT: MISMATCH (HIGH)
recommendation: add an educational content layer (guides, docs, FAQ depth)
SERP result classification (heuristic - titles and URLs only):
1 Blog https://developers.google.com/search/docs/appearance/ai-feat
2 Blog https://search.google/ways-to-search/ai-overviews/
3 Blog https://www.reddit.com/r/GEO_optimization/comments/1tujdzj/h
4 Tool https://www.seo.com/tools/ai-overview-checker/
5 Blog https://ahrefs.com/blog/how-to-track-ai-overviews/
6 Tool https://seranking.com/ai-overviews-tracker.html
7 Blog https://www.semrush.com/blog/semrush-ai-overview-research/
8 Blog https://www.youtube.com/watch?v=kqjFC5L_2jo
types: Blog 6 | Tool 2
AI Overview: present (2444 chars), citations incl. api.dataforseo.com, encrypted-tbn0.gstatic.com, support.google.com, www.botify.com, www.youtube.com
PAA questions (4):
- How do I find my AI Overview?
- Can you turn off AI Overview?
- Why isn't the Google AI Overview showing up on my account?
- How do I turn on AI in Google Search?
Related searches (8): How to check ai overview free | AI Overview Search | AI Overview Google | AI Overview website | AI overview app | AI Overview Google turn on | AI Overview Download | How to use Google AI overview
Ads on capture: 0
SXO gap score (separate from any technical-health score):
Page type 5/15 Tool mismatch severity HIGH
Content depth 13/15 1011 words, 12 H2 headings
UX signals 10/15 viewport meta, one h1
Schema markup 15/15 WebApplication present, valid JSON-LD @context, FAQPage block
Media richness 10/15 4 images, alt coverage 4/4
Authority signals 15/15 Organization/WebSite schema, author entity, social proof phrases, years / numbers
Freshness 0/10 no date signals (tool page may legitimately not need one - judgment call)
TOTAL: 68/100Now the page did not change, but the SERP rewarded Blog (6 of 8) and the Tool result at #4 is the very same seo.com checker that ranked at #1 for the commercial term. Same competitor holding both shapes. Our score drops from 78 to 68 entirely on the Page type line: 15/15 → 5/15. Not because anything on the page got worse, but because the keyword changed what the page is being compared with.
Also new: the AI Overview is present with 2444 characters this time. Watch the citation list. api.dataforseo.com is the SERP API's own domain and encrypted-tbn0.gstatic.com is Google's CDN for image thumbnails; both are capture artifacts, not sources the AI Overview reads from. The three real ones are support.google.com, www.botify.com, and www.youtube.com. Treating artifacts as citations would be a small, plausible-sounding error, and the recommendation is to eyeball that line in every run.
The PAA set here is the diagnosis in four lines: "How do I find my AI Overview?", "Can you turn off AI Overview?", "Why isn't the Google AI Overview showing up on my account?", "How do I turn on AI in Google Search?". Finding, turning on, turning off. Those are help-center questions. Google's own documentation and Reddit answer them. A tool page that assumes you already found one does not.
The control run
Same day, same capture window, the sister tool page (auspia.ai/tools/google-index-checker) against its own keyword "google index checker": Tool 70%, STRONG consensus, VERDICT ALIGNED, 77/100, and its Freshness line scored 4/10 with "visible 'updated' text", where the AI Overview page scored 0. So the two-fingered result above is not the page; it is the keyword. Which is the honest way to phrase the finding: run two is not "our page is broken", it is "this keyword wants something our page is not".
Reading the verdicts honestly
Three things take practice.
ALIGNED is not a score of 100. It is a statement of that one dimension: the page type matches. 78/100 means the SERP did not see a problem on the type axis, and 22 points were left on freshness (no date signals, and the script itself flags that an evergreen tool page may legitimately not need one: a judgment call, not a defect), UX, and content depth. ALIGNED pages still compete on depth and authority against 8 tools; that competition is what technical audits are for.
MISMATCH HIGH with "add an educational content layer" does not mean "rewrite the tool page to be an article". It means the SERP rewards informational content, and a page that is a tool can carry that layer. For this page, that means a guide at /blog/how-to-check-ai-overviews-style URLs answering exactly the PAA questions: find it, turn it on, turn it off, test it, all linking into the tool. One page type, one keyword, no cannibalization argument: the guide targets informational, the tool page targets commercial.
And the order of operations from the skill's own rubric: fix the weak persona first, then the weak dimension across personas, then the mismatch. For this page the mismatch is already known (we have the scan), so the weak persona is what comes next.
Persona cards, derived from the same signals
After the scan, the skill's second job: build 3-5 personas, each traceable to a signal the scan printed, score the page from each persona's perspective on 25 points per dimension, and rank fixes. These cards are from the two runs above. The signals in the left column are verbatim from the scan; the scores are judgment, and I am labeling them as such.
Persona | Signal it traces to | Relevance | Clarity | Trust | Action | Total |
|---|---|---|---|---|---|---|
Quick-Answer Seeker | PAA "How to check AI Overview?" (run 1), "How do I find my AI Overview?" (run 2) | 14 | 12 | 17 | 16 | 59 |
Free-Tool Comparison Shopper | Related searches "Best ai overview checker", "Ai overview checker free", "Ai overview checker online" | 23 | 22 | 20 | 21 | 86 |
AI-Informed Reader | AI Overview element present, 2444 chars, cites support.google.com, botify.com, youtube.com | 18 | 16 | 21 | 15 | 70 |
One user story, in the skill's required format:
As a Quick-Answer Seeker, I want to check whether my page shows up in AI Overviews, because the client asks me every week and I still cannot tell if my last three changes did anything, but I'm blocked by the fact that the page opens with a form, not with the words "type your keyword below and it will list the domains".
That barrier is specific enough to be a fix: the above-fold area answers the "why should I trust your search box" question before the search box. And the Free-Tool Comparison Shopper scoring 86 tells you the commercial term is being served. The 59 is the reader you are losing first.
Where this still stops
- The SERP classification is based on titles and URLs only. It finds the
consensus; it does not describe each competitor's page. If a result's title is misleading, the classification is too, and it is visible in the report.
- One capture is one moment. Google's SERP moves; the verdict is
day-stamped, not eternal.
- An AI Overview JSON item can exist without its text (async render). The
skill says "treat as present" rather than assuming absence.
- Fragmented SERPs (< 40% dominant) are an opportunity, not a mismatch,
and skipping severity is the honest move there.
- SXO is not technical health. Run the page-audit skill for on-page
scoring, the technical-audit skill for the crawl side.
- A paid SERP API costs a bit per query (the two captures for this article
were about $0.02). The TSV path is free: copy titles and URLs from a search you ran yourself.
- Persona cards, user stories, and wireframes are judgment. They are model
output constrained to cite extracted signals, never measurements.
- JS-rendered pages may need the rendered HTML saved locally and passed as
a file; the fetch path can also be blocked by shields and bot checks.
- Nothing scrapes Google. The SERP file is a capture; if you cannot get
one with at least 5 organic results, the skill says so instead of inventing a consensus.
Troubleshooting
Failure | What it means | Fix |
|---|---|---|
| Bot check on the target page | Save the page's rendered HTML locally and pass the file path instead of the URL |
| Transient SSL/network error | Run it again; it usually succeeds on retry |
| The JSON isn't a valid serp/google/organic response | Re-export the raw response; check the API status message |
"only N organic results ... limited confidence" | Fewer than 5 organics | Find a denser capture or run with less certainty and say so |
Empty SERP prints the refuse-to-score message | No organic data | The gate worked; feed it a real capture |
Empty or wrong output from a TSV | Missing tab, or title/URL in the wrong column | One |
| Script runtime missing |
|
Pasted script has literal backslash-n or backslash-x sequences | The shell interpreted them and corrupted the file | Save the |
AI Overview line says async-render | The capture has the element without text | Trust the existence, and fetch a fresh capture if you need the citations |
PASTE-TO-CODEX
Want Codex to do the installation? Copy everything from the line above this paragraph to the end of the article, paste it into Codex, and say this:
Install the two files in this article as a Codex skill.
Create ~/.codex/skills/codex-seo-sxo/SKILL.md verbatim from the
~~~~markdown block and scripts/sx_scan.py verbatim from the ~~~~python
block, keeping the file layout shown in the install section. Also create
scripts/smoke.tsv from the sample TSV block in the install section. Then
run:
python3 ~/.codex/skills/codex-seo-sxo/scripts/sx_scan.py https://auspia.ai/tools/ai_overview ~/.codex/skills/codex-seo-sxo/scripts/smoke.tsv
Report the paths you created and the output. If python3 is missing, report
that instead of stopping. Do not install or modify any other files.If the smoke test prints SERP consensus: Tool and VERDICT: ALIGNED, the skill is live. Your first real run is against your own page: save your keyword's organic results as JSON or TSV, point the script at both, and wait for the verdict and the 7-dimension score. Read the verdict before you rewrite anything, including this sentence, because the verdict tells you whether rewriting is even the right move.
FAQ
Do I need a paid SERP API? No. The TSV path is free: titles and URLs from a search you ran. JSON just adds PAA, related searches, and AI Overview text, and that extra material is what personas are built from.
What does "page type" mean exactly? Landing, Blog, Product, Hybrid, Service, Comparison, Local, or Tool. The taxonomy above lists the primary signals and the typical content structure for each. A page gets exactly one type, resolved by the priority order.
My page says ALIGNED but it is still position 10. Why? The type matches, so the problem is elsewhere: depth, authority, technical health, or competition. Run the page-audit and backlinks skills. The SXO verdict is the match axis only, and the skill says so to stop you from blaming the wrong thing.
What if my keyword's SERP is one big AI Overview with three organic results? The script prints the limited-confidence note and does not score a consensus. Then answer with the ad themes and PAA questions, and be honest that the organic consensus is thin. AI Overview text with fewer organics is a signal to read, not a number to manufacture.
Why does an article that teaches SEO tell me when not to write an article? Because the query type decides that, not a preference for content. Getting a wrong-type page to rank is the most expensive rewrite cycle in SEO: it compounds. This skill exists to check the premise.
Can the skill tell me my persona scores by itself? No. The script extracts PAA, related searches, ads, and AI Overview citations, and the first run prints them. Persona cards, user stories, and IST/SOLL wireframes are the model's job in the same run, with the rule that every persona must trace to a printed signal. That is the part where the skill gets less mechanical and more useful, and it is also the part you should check by hand.
This is post 16 of the Codex SEO Skills series: a 20-part set that reads Google's own decision structure into skills any Codex install can run. Each part is a self-contained skill; install the preflight skill (codex-seo-ready) first, then this one, and the rest as your workflow needs them. Skills in the same family share the `~/.codex/skills/` layout, and in this series two neighbors are the closest to SXO: the page audit (codex-seo-page-audit) scores on-page health, and the GEO skill (codex-seo-geo) covers the AI surfaces that increasingly ARE the SERP.
Author: Freya Collins, AI Answer UX Researcher, 900+ Answer Blocks Studied. Freya writes about how people actually read answer surfaces, and why a page's shape decides whether it gets read at all.
Based on the open-source [claude-seo](https://github.com/AgriciDaniel/claude-seo) project (MIT license, AgriciDaniel). This article customizes the SXO skill for the Codex runtime and keeps the original methodology: the eight-type page taxonomy, the SERP consensus thresholds, the mismatch severity table, the seven-dimension gap score, and the signal-to-persona rule. The scan script is written from scratch for this article.



