Short answer: what semantic SEO is
Semantic SEO is optimizing a page for a topic and the web of ideas around it, instead of a single keyword phrase. In 2026 the practical difference is stark: search engines and AI answer systems want clear, extractable answers about entities (people, places, products, concepts) and how they relate to each other. Structured data helps confirm what you are, but it is not what gets you cited. Visible, answer-first writing is.
Here is the short version before the details: pick one topic, map every real question a reader brings to it, answer those questions directly and up front, mark who you are and what you know, connect related pages by topic, and use schema to state your identity, not as a shortcut to rank. That is semantic SEO, and you can do all of it by hand or hand most of it to Codex with the skill at the end of this guide.
Watch: semantic SEO in 76 seconds
Prefer watching to reading? This 76-second explainer walks through the same five-part system with the 2026 evidence: what schema does and does not do for citations, why AI answer systems reward extractable structure, and where the ~38% citation-ranking gap leaves an opening for new pages.
Why semantic SEO matters more now
Two things changed between the old "keyword" playbook and 2026.
First, AI answers became a normal way people search. Google's AI Overviews now appear on roughly half of tracked queries, and AI Mode plus ChatGPT, Perplexity, and Gemini answer questions directly. These systems do not match keyword strings. They retrieve passages that clearly answer a question and cite the page they came from.
Second, being cited stopped being the same thing as ranking #1. In a 2026 analysis of 863,000 keywords, only about 38% of pages cited by Google's AI Overviews also ranked in the top 10, down from 76% a year earlier. A page on page two can appear in an AI answer if its structure and content are extractable. That is a real opening for beginners: a well-built page on a clear topic can be pulled into an AI answer without winning a classic ranking.
Freshness counts too. Ahrefs reports that around 65% of AI crawler hits land on content published within the last year. Old pages that were never updated are quietly losing their seat in AI answers.
The 2026 reality check: schema will not get you cited
This is the single most important correction to the old semantic SEO advice. For years, guides (including the one this article refreshes) told you to add structured data to get an edge. In 2026, the controlled evidence says that edge does not exist.
In May 2026, Ahrefs tracked 1,885 pages that added JSON-LD schema and compared them with nearly 4,000 matched control pages that never added it. Across Google AI Overviews, AI Mode, and ChatGPT:
Platform | Change in AI citations after adding schema |
|---|---|
Google AI Overviews | −4.6% |
Google AI Mode | +2.4% (statistically indistinguishable from zero) |
ChatGPT | +2.2% (statistically indistinguishable from zero) |
A companion experiment by searchVIU fetched pages with five AI systems (ChatGPT, Claude, Gemini, Perplexity, and Google AI Mode) and found none of them read hidden JSON-LD, microdata, or RDFa. They extract the visible HTML, the same text a human sees.
So why do most AI-cited pages still have schema? Correlation. Well-maintained sites tend to have both better content and proper schema. The schema rides along with the signals that actually matter. Google's own guidance says AI Overviews and AI Mode require no special schema, no llms.txt files, and no AI-specific optimizations, and it calls "overfocusing on structured data as a shortcut" a mistake. Google also deprecated FAQ rich results in May 2026.
That does not mean schema is useless. It is useful as identity infrastructure: Organization, Person, and Article markup with consistent sameAs links tell search engines which entity you are, which matters for disambiguation. Just stop expecting it to move citations.
Your five-part semantic SEO system
This is the workflow. Run the parts in order, by hand first. Every part includes what to do, what good looks like, and a quick check.

Figure: The five parts of semantic SEO. Each part answers a different layer of the same question: what does this page mean, and who is it for?
Part 1: Research the topic, not just keywords
Start with one topic and widen it into the full set of questions people actually ask.
- Write down your topic as an entity: "running shoes for beginners," not "best running shoes."
- Open Google and note the suggestions, the People Also Ask block, and the related searches at the bottom. Those are the questions real people type.
- Add synonyms and related concepts. For running shoes: cushioning, arch support, trail vs road, pronation, drop, shoe width, what brands to avoid.
- Group everything by intent: questions that are learning ("what is drop in running shoes"), comparing ("trail vs road shoes"), or choosing ("best shoes for flat feet").
- Map which of those questions your page can answer well, and which belong on other pages of your site.
What good looks like: a list of 10–20 related questions grouped by intent, with the ones your page will answer clearly marked.
Quality check: can you answer every kept question in two sentences from your own knowledge? If not, either learn it or drop it. Recovery: if your list feels thin, search the topic on Wikipedia and mine its subheadings and references for entity names you forgot.
Part 2: Write answer-first, entity-clear content
Now write the page so both a human and an AI system can find the answer in seconds.
- Lead with the answer. Put the direct answer in the first paragraph. This is called bottom-line-up-front (BLUF) writing, and it is the highest-signal formatting change you can make. Do not make readers scroll through history to reach the point.
- State your entity clearly. If you are "SoleMate," a shoe shop in Portland, say so in plain words near the top: who you are, what you offer, where you operate. An AI system cannot attribute a page to an entity it cannot name.
- Show evidence a human wrote it. A named author, a note on how you tested the shoes, photos from your own floor, original numbers. AI systems cite what they cannot easily synthesize themselves: first-hand experience and original data.
- Keep it fresh. Update the page at least once a year. Add a "last updated" date where the theme allows.
A weak opening: "In the footwear industry, many factors influence the consumer decision-making process when selecting athletic shoes."
A strong opening: "Most beginners buy the wrong running shoes because they pick a brand first. Start with your foot shape, your arch, and the surface you run on. Then pick a shoe that fits those three facts. Here is how to do that in five minutes."

Figure: The same topic, two openings. The version on the right is the one a reader remembers and an AI system can quote.
Quality check: cover the answer with the question visible, and any reader should be able to say what the page is about in one sentence. Recovery: if the page is a wall of text, cut every paragraph that is not answering one of your Part 1 questions.
Part 3: Structure for extraction
AI systems pull passages, not whole pages. Make the passages easy to pull.
- Use question-shaped headings. "What is a running shoe drop?" is far more citable than "Drop."
- Keep one idea per block: short paragraphs, lists, and tables. A table comparing trail vs road shoes is extractable; a paragraph that squishes the same facts together is not.
- Use semantic HTML: real heading levels (one H1, then H2/H3), real lists, real
<blockquote>for quotes, real<table>for data. These are structural hints, not decoration. - Never let an important fact live only in JSON-LD. If you strip all markup, the fact must still be on the page in visible text. That is the single rule that predicts AI extractability.
Quality check: copy the visible text of your page into a plain text editor. Can a stranger still find the answer? Recovery: if a key fact appears only in a schema block or an image, rewrite it into a sentence or a table row.
Part 4: Link by topic relationship
Internal links tell search engines how your pages relate. In semantic SEO, you link by topic, not by "here are our other posts."
- Build a pillar page for the big topic and link it to cluster pages for the subtopics. Your "running shoes" pillar links to "how to choose running shoes," "trail vs road shoes," "best shoes for flat feet."
- Use descriptive anchor text. "Our guide to trail vs road shoes" beats "click here." The anchor text is a relationship statement.
- Link both directions: cluster pages link back up to the pillar, and pillars link down to clusters.
What good looks like: every cluster page can be reached from the pillar in two clicks, and the anchor text on each link describes the relationship.
Quality check: remove all links and check whether each page still reads well. Links should feel like the site has structure, not like a link farm. Recovery: if a cluster page has no incoming topic links, that subtopic probably deserves its own page, or should be merged into the pillar.
Part 5: Add schema as identity infrastructure
Now that the visible content is strong, add schema to confirm identity, not to chase citations.
- Article schema on every article, with the author and publisher named.
- Organization schema on your site, with consistent
sameAslinks to your LinkedIn, Crunchbase, and Wikidata pages if they exist. - Person schema on author pages.
- Keep the markup consistent with the visible text. The name in schema should match the name on the page.
That is it. Skip FAQ and HowTo schema chasing the old rich-result era; those features are gone or retired. Validate what you add with Google's Rich Results Test.
Quality check: your schema says the same name, address, and facts as your visible content, and it validates. Recovery: if Google cannot tell you apart from a company with a similar name, that is a sameAs and consistency problem, not a schema-amount problem.
Hand the whole workflow to Codex
Once you understand the five parts, you can automate most of the research and the page check with Codex, OpenAI's coding agent. The skill below is a complete, copy-paste file. It reads your topic, gathers related questions from the sources you approve, clusters them by intent, and reports what it could and could not verify. It never invents numbers and never asks for secrets.
How to use it as a beginner
- Install the free Codex CLI and open it in a folder you control.
- Create a folder named
.codex/skills/semantic-seo/inside your project and save the block below asSKILL.md. - Tell Codex: "Use the semantic-seo skill for the topic: running shoes for beginners, US market, English."
- If you have a paid data source (Ahrefs API, Semrush API, DataForSEO) or a Google Search Console export, connect it and say so. Otherwise Codex will use public surfaces only.
- Read the report it produces. Approve nothing that changes your site until a human checks the "gaps" and "limits" sections.
# Semantic SEO Research Skill
Runs a topic-based semantic SEO research pass for one page or topic.
## Inputs
- topic or page URL (required)
- target market/country (e.g., "US")
- target language (e.g., "en")
- optional: authorized data sources (Ahrefs API, Semrush API, DataForSEO, or a Google Search Console export)
## Prerequisites
- The user has configured any paid data source and approves its use.
- No API keys, cookies, passwords, or tokens are requested, read, or printed.
- If no paid source is configured, use only public surfaces: Google suggestions, People Also Ask, related searches, and Wikipedia.
## Steps
1. Define the topic and its core entity (who/what the page is about).
2. List the main questions a reader brings to the topic.
3. Collect related terms, synonyms, and subtopics from the approved sources.
4. Group them into intent clusters: learn / compare / choose / act.
5. For each cluster, identify the entity and its relationships.
6. Mark each item as confirmed, estimated, or unavailable, with its source.
7. Output an opportunity list with notes for human review.
## Output
A Markdown report containing: the topic, an entity map, intent clusters, per-cluster questions, data provenance (provider, report, retrieval date), gaps, and next actions.
## Limits
- Never invent search volume, difficulty, CPC, or SERP results.
- Never include private data or credentials in any output.
- Never promise rankings, traffic, or AI citations.Safety rules for everything you hand to Codex: keep your API keys out of the chat, approve each edit before it touches your site, and treat every metric it prints as "reported by the source," never as guaranteed truth.
Beginner mistakes that break semantic SEO
- Forcing the exact keyword. Semantic SEO rewards topic coverage. A page that says "best running shoes" forty times reads as spam to humans and adds nothing for AI.
- Hiding facts in schema. If the answer only exists in JSON-LD, AI systems fetching your page will miss it. Visible text first, always.
- Treating schema as a rank lever. It confirms identity; it does not win citations. See the Ahrefs study above.
- Ignoring intent. Answering a "which is better" question with a definition page helps no one. Match the question type.
- Never updating. If your page is older than a year, treat it as stale. AI systems skew toward recent content.
- Writing "AI bait." Chunked, answer-first formatting with nothing real behind it gets cited for the wrong reasons and hurts trust. Write for a reader who needs the answer.
Verify your work: a 60-minute check
Run this once a quarter on your most important pages:
- [ ] Can a stranger state what the page is about after one sentence?
- [ ] Is the direct answer in the first paragraph?
- [ ] Are headings question-shaped and real H2/H3 levels?
- [ ] Would the key facts survive if all markup were stripped?
- [ ] Are related pages linked with descriptive anchor text?
- [ ] Does Organization/Article/Person schema match the visible text and validate?
- [ ] Was the page updated in the last 12 months?
- [ ] Does the page read as written by a human with experience, not a script?
If you answer "no" to more than two, start with Part 1 again on that page.
FAQ
Is semantic SEO the same as entity SEO? Entity SEO is the core of semantic SEO. Semantic SEO is the broader practice: researching the topic, answering questions, linking by concept, and structuring content. Entity SEO focuses specifically on making sure search engines know exactly which entity your page is about and how it connects to other known entities.
Do I need to be an SEO expert to do this? No. Every part of the five-part system is a writing and structure habit, not a technical specialty. The hardest part is honestly listing what your reader wants to know, and you can do that without any tool.
Does schema still matter for rich results? Some types do. Product, Review, Event, and Video schema still support rich result features. But FAQ rich results were deprecated in 2026, and schema does not drive AI citations. Keep it as identity and feature infrastructure.
Will this guarantee an AI Overview or a top ranking? No. Nobody can guarantee either. The evidence shows answer-first, entity-clear content raises your chances of being retrieved and cited, but search results are competitive and noisy. Treat semantic SEO as making your page the best possible answer, not a guarantee.
What is the quickest win for a total beginner? Rewrite the opening paragraph so the direct answer comes first, and add a "last updated" date. Both take ten minutes and both are exactly what AI retrieval systems reward.
Author: Clara Bennett, 10-Year Content Strategy Practitioner at Auspia. Clara writes about editorial systems, topic maps, repeatable content operations, and how SEO and GEO workflows actually run inside a growth team.












