How to Get Cited by AI Search: Fix the 4 Factors That Actually Matter

Key takeaways

A viral post claims schema markup has zero effect on AI citations, based on a study that never tested schema at all. Here is the real study's 4-factor priority order, how to audit your page against it, and what to fix first.

What you will finish with

A page-level audit against the four factors a controlled 252,000-trial study found decisive for AI citation, plus a fix order for the seven secondary factors that matter once those four are in place. You will also know, with sources, why a popular claim that "schema markup has zero effect on AI citations" is not what that study actually found, so you do not waste a sprint chasing the wrong fix.

Who this is for: anyone who owns page-level SEO or GEO for a site trying to get cited in ChatGPT, Gemini, Perplexity, Claude, or Google's AI answers.

Time required: 30-45 minutes per page for the audit and first-pass fixes.

Done when: the page you are auditing passes all four gatekeeper checks, and you have a written list of which secondary factors still need work.

Before you start: what this checklist is based on, and what it corrects

A post circulating in SEO/GEO circles claims a SIGIR study ran 252,000 trials across six LLMs and found schema and JSON-LD have "zero measured effect" on AI citations, while price, freshness, specs, and depth move the needle hard.

The study is real. The trial count is right. The odds ratios in that post are close to real numbers. The part that does not hold up: that study never tested schema markup. It is not that schema "measured zero" — schema was never on the list of things measured.

The paper is "What Gets Cited: Competitive GEO in AI Answer Engines" (Vishwakarma, Kumar, Jamidar, Sprinklr; SIGIR '26, Melbourne). Design: inject exactly two candidate sources into six LLMs (Gemini-2.5-Flash, Claude-3.5-Sonnet, Kimi-K2-Thinking, three GPT-5 variants), each pair differing in exactly one of 18 content factors, and record which source gets cited first. The 18 factors span content match, completeness, trustworthiness, readability, competitive standing, and freshness. Structured data format is absent from all 18 — confirmed directly against the paper's own taxonomy table.

If you want the real answer on schema, it comes from a separate source: Ahrefs' May 2026 controlled study of 1,885 pages that added JSON-LD, matched against 4,000 control pages. Result: no meaningful citation lift on Google AI Overviews (-4.6%, a small but statistically significant decline), Google AI Mode (+2.4%, not significant), or ChatGPT (+2.2%, not significant). A separate searchVIU test found that ChatGPT, Claude, Perplexity, Gemini, and Google AI Mode do not read JSON-LD, hidden Microdata, or hidden RDFa during live retrieval at all.

Keep both facts straight going into this audit: the SIGIR study tells you what to prioritize on the page. The Ahrefs study tells you schema is not one of those priorities for citation volume specifically, even though it still matters for rich results and entity clarity elsewhere.

Step 1: Run the gatekeeper audit

Check your page against these four factors. The SIGIR study found all four unanimous across all six models, with very large effects (odds ratio over 100 in most models). Failing any single one can eliminate citation odds regardless of how strong the rest of the page is.

Gatekeeper

Check

Odds ratio range (across the 6 models tested)

Topic match

Does the page directly answer the query it is meant to rank for, not something adjacent?

Effectively decisive in 5 of 6 models

Price stated

Is a specific price visible on the page, not gated behind "Contact us"?

6.26 to over 10,000

Recent timestamp

Is a current published or last-reviewed date visible and accurate?

14.4 to over 10,000

List position

Does this page tend to rank first among competing candidates for the query, not second?

1,795 to over 10,000

Action: open the page next to the target query. Read the first 150 words and ask if they answer the query directly. Search for the price on the page; if it requires a form or a call, that is a fail. Check the visible date against when the content was actually last updated. Check your current rank or citation-share for the query.

Expected output: a pass/fail for each of the four rows.

Quality check: if you fail topic match, stop here — none of the later steps will fix a page that answers the wrong question.

Recovery path: price and timestamp fails are same-day fixes (Step 2). A list-position fail is not a content fix; it means the underlying SEO fundamentals need work before this page is competitive for citation at all, and no amount of on-page polish here will substitute for that.

Flowchart showing the four gatekeeper checks in sequence, with a stop-and-fix branch for any failure
Any single gatekeeper failure can eliminate citation odds — fix it before moving on.

Step 2: Fix any gatekeeper failures first

Do these before touching anything else. They are the highest-leverage, lowest-effort items on this list.

If price is missing: add the actual number. If pricing varies, publish a real starting price or a clear range instead of a contact form. This was one of the four gatekeepers in the study.

If the timestamp is stale or missing: update the published or last-reviewed date, but only after you have actually revised the content. A date change with no real update is a freshness signal you cannot back up if anyone checks.

If topic match is weak: rewrite the opening 150 words to directly answer the query's task, not a related topic. Do not bolt a new intro onto old structure; replace it.

Quality check: re-read the page as if you were the query. Would a stranger get the specific answer in the first paragraph?

Recovery path: if a genuine price cannot be published (custom quotes, enterprise deals), publish a representative starting price or typical range instead of leaving it blank — the study measured "price not mentioned" as the failure state, not "price varies."

Step 3: Work through the secondary differentiators

Once the four gatekeepers pass, these seven factors provide secondary differentiation, with odds ratios ranging from about 2 to 243 depending on the model. Fix them in this order:

  1. Missing specifications. Add the technical specs a buyer would actually compare (dimensions, materials, capacity, compatibility). Odds ratio 8.63 to 243 across models — one of the largest secondary effects.
  2. No comparisons. Add a direct comparison against at least one named alternative. Odds ratio 1.61 to 7.45.
  3. Hedged language. Replace "might," "could," "possibly" with direct, confident statements the page can actually support.
  4. Unsupported claims. Attach evidence to claims: test results, certifications, named methodology.
  5. Internal contradictions. Check that numbers and claims agree across sections of the same page; this is easy to introduce during edits and easy to miss on review.
  6. Keyword gaps. Confirm the page actually contains the specific terms the target query uses, not just synonyms.
  7. Weaker value proposition. Sharpen the specific benefit claim against what a competing page offers.

Action: go down the list in order and fix what applies. Not every page will have all seven issues.

Expected output: a page where every claim is either specific, comparative, or evidenced.

Quality check: would a skeptical reader find a single vague or unsupported sentence left on the page?

Recovery path: if you are short on time, stop after items 1 and 2 (specs and comparisons) — they carry the largest odds ratios of the seven.

Bar chart ranking the seven secondary AI-citation factors by odds-ratio range, from missing specifications at the top to weaker value proposition at the bottom
Once the four gatekeepers pass, work down this list in order.

Step 4: Skip formatting-only changes

The study explicitly found no consistent effect from restructuring paragraphs into bullet points, adding subheadings for scannability, or reorganizing scattered information, because the tested models parsed content regardless of visual structure. If your page passes the checks above, do not spend a sprint on formatting polish expecting a citation lift from it. Format for your human readers, not for a citation gain the data does not support.

Step 5: Leave schema where it belongs

Keep schema markup technically valid, because it still does real work: it helps Google and Bing parse rich results, and it gives search systems an unambiguous read on what a page and entity are. What it does not do, per the best available controlled evidence, is move AI citation odds. Do not move schema work ahead of Steps 1-3 on the theory that it is a citation lever — no controlled study supports that, and the one study that gets cited for it never tested schema at all. If you want the full breakdown of what the Ahrefs schema study actually measured, see our 2026 semantic SEO guide.

Verify the finished result

Re-run the Step 1 audit on the page after your fixes. All four gatekeepers should now pass. Spot-check two or three of the secondary factors from Step 3 that you fixed. If you have access to an AI visibility or citation-tracking tool such as the AI Search Visibility Checker, note the page's current citation rate for its target query as a baseline before moving to the next page, so you can tell later whether the changes moved anything.

Maintain the result

Recheck the timestamp gatekeeper every time the page gets a substantive update, not on a fixed calendar. Re-audit pages that lose citation share, starting with the four gatekeepers before assuming a secondary factor slipped.

FAQ

Did the SIGIR study test schema markup or JSON-LD? No. Its 18-factor taxonomy covers content match, completeness, trustworthiness, readability, competitive standing, freshness, and list position. Structured data format is not among the categories or the factors tested — confirmed by checking the paper's taxonomy table directly.

Does schema markup help AI citations at all? The best available controlled evidence — Ahrefs' May 2026 study of 1,885 pages against 4,000 matched controls — found no meaningful citation lift from adding JSON-LD on Google AI Overviews, Google AI Mode, or ChatGPT. A separate test found several major AI models do not read structured data during live retrieval. That does not make schema worthless elsewhere; it is just not a citation-volume lever.

Where did the specific numbers in the viral post (6.3, 14, 4 to 8.6) come from? They resemble real odds ratios in the SIGIR paper's results table, but they belong to price, recency, and specs/comparisons factors, not schema. The paper reports figures per model rather than as one blended average, and ranges vary substantially across the six LLMs tested.

What should I fix first if I only have an hour? The four gatekeepers: topic match, price, timestamp, and list position. Failing any one of them can eliminate citation odds regardless of everything else on the page.

Should I stop working on schema entirely? No. Keep it valid for rich results and entity clarity. Just do not prioritize new schema work over the gatekeeper fixes if AI citation volume is the actual goal.

Author: Bennett Hayes, Applied GEO Analyst Across 400+ Implementation Reviews at Auspia. Bennett writes about practical GEO execution, audits, and implementation notes grounded in what controlled studies actually measured.

Explore this topic

Keep following the same growth thread