Updated July 29, 2026
10 min read
Tactics

Which Schema Types Actually Help With AI Answers?

What Google Still Supports, What It Quietly Dropped, and What the Citation Evidence Shows

Short answer

No schema type is required for AI answers, and Google states there is no special markup you need to add. Article, Organization, BreadcrumbList and Dataset remain supported Google features worth keeping, while FAQPage, HowTo and ClaimReview no longer produce rich results and should not carry a strategy.

GB
Written byGabor BartaCo-founder, Oppira

Gabor leads product and content at Oppira. He has spent over a decade building tools and writing about competitive intelligence, social media analytics, and growth marketing for B2B SaaS companies.

Published July 29, 2026

1. Does structured data help a page appear in AI answers?

Not as an entry requirement. Google states plainly that structured data is not required for generative AI search and that there is no special schema.org markup you need to add for its AI features.

Structured data
Machine-readable markup, usually JSON-LD in the page head, that describes what a page is about using the schema.org vocabulary.
Also known as: Schema markup, JSON-LD, Rich result markup

The eligibility rule Google publishes for AI Overviews and AI Mode mentions no markup at all. A page has to be indexed and eligible to be shown in Google Search with a snippet, and Google adds that there are no additional requirements and no other special optimisations necessary. Markup is not one of the gates.

None of the other engines contradict that. OpenAI, Anthropic, Perplexity and Microsoft document their crawlers and their retrieval behaviour, and none of them document a schema type that changes source selection. Any claim that a specific markup type is weighted during selection is an inference nobody has published evidence for.

The right frame is that markup is cheap insurance, not a lever. It costs almost nothing to emit correctly, it powers a handful of surviving Google rich results, and it is semantically honest. It is not the thing standing between a page and a citation.

2. Which structured data types are worth keeping in 2026?

Article, BreadcrumbList, Organization with sameAs, ImageObject and SoftwareApplication are all still supported Google features. Person on the author and Article.citation are worth adding even though neither produces a rich result.

Schema types worth emitting in 2026, with the reason each one earns its place
TypeVerdictWhy
Article or TechArticleKeepSupported Google feature. Carries dateModified, author, citation, about and image, all of which should also be visible in the page text
BreadcrumbListKeepSupported Google feature and the cheapest markup on this list to maintain correctly
Organization with sameAsKeep and extendSupported Google feature. sameAs pointing at Wikidata, LinkedIn, Crunchbase, G2 and Trustpilot is the main entity-linking mechanism a site owner has
Person on the authorAddNo rich result. E-E-A-T signals were the second strongest factor in the Semrush analysis of 304,805 cited URLs, and a named accountable author is the visible half of that
Article.citationAddNo rich result. The GEO paper at ACM SIGKDD 2024 measured a 28% improvement on its visibility metric from citing sources, and the markup mirrors what is already on the page
DatasetAdd if you publish measured aggregatesSupported Google feature through Dataset Search. No published evidence that it affects AI citation, and almost no competitor emits it
ImageObject inside Article.imageAddGoogle explicitly recommends high-quality images and states they can surface beyond web page links
DefinedTerm inside a DefinedTermSetKeep on glossary pagesNo rich result. Cheap, semantically true, and it reinforces the term cluster a site is trying to own
SoftwareApplication or WebApplicationKeep on tool pagesSupported Google feature for software listings, and accurate for a calculator or an app page
Schema types worth emitting in 2026, with the reason each one earns its placeEvery item in this table describes something that is genuinely true about the page. That is the test worth applying: if the markup asserts something the page does not contain, it is a liability rather than an asset.

3. Which structured data types should I stop investing in?

FAQPage, HowTo, ClaimReview and Speakable. Google stopped showing FAQ rich results on 7 May 2026 and deprecated HowTo in September 2023, and neither type is in the structured data search gallery any more.

Schema types and files that no longer earn the effort, and what to do with each
TypeWhat happenedWhat to do
FAQPageGoogle stopped showing FAQ rich results on 7 May 2026, dropped the search appearance and Rich Results Test support in June 2026, and removes Search Console API support in August 2026Leave existing markup in place and invest nothing further
HowToDeprecated as a rich result in September 2023 and removed from the structured data search galleryEmit a semantic ordered list instead, because the HTML is what a crawler parses
ClaimReviewSearch Console support is being removed, and the programme is limited to approved fact-check publishersDo not add
SpeakableRestricted to news publishers, with no mechanism by which another site would benefitDo not add
llms.txtNot read by any engine. Google states Search does not use AI text files, and Ahrefs found 97% of files across 137,210 domains received zero requests in May 2026Keep an existing file, do not expand it
Any AI-specific markup vocabularyNone exists. Google states there is no special schema.org structured data that you need to add for its AI featuresDo not add
Schema types and files that no longer earn the effort, and what to do with eachRemoving FAQPage markup that already exists is not worth the deployment. Adding it to fifty new pages because a checklist says to is an afternoon spent on a feature Google has switched off.

The FAQ case is worth understanding rather than just accepting, because it explains the whole category. What correlated with citation in the Semrush data was the visible question and answer text on the page, at plus 25.5%, not the JSON-LD describing it. The markup was always a wrapper around the thing that mattered, and when Google switched the rich result off, the wrapper lost its last measurable job.

4. Are pages with schema markup cited more often?

Observational datasets show a mild positive association, and the one controlled test found formatting and structure negligible. The likeliest explanation is that competent publishers emit markup and also write better pages.

The evidence genuinely conflicts, so it is worth laying out both sides. Semrush, across 11,882 prompts and 304,805 cited URLs, found structured data elements associated with citation at plus 21.6%. The GEO-16 analysis by Kumar and Palkhouski, covering 1,702 citations across 16 B2B SaaS verticals, reported structured data at r=0.63 and semantic HTML at r=0.65.

Both are observational. Against them sits the SIGIR 2026 paired test by Vishwakarma, Kumar and Jamidar, which changed one factor at a time across 252,000 trials and six models, and found formatting and structure changes negligible. And Citera, analysing roughly 350,000 B2B SaaS articles, found schema markup present on 69 to 72% of articles at every ranking position, with no gradient whatsoever.

A confound explains all four results at once. Sites that emit clean markup tend to be sites with a competent publishing setup, which also produces clearer writing, better dates and more original content. The correlation is real and the causal arrow points at the publisher, not the markup.

5. Which schema types are underused and worth adding?

Dataset, Person on the author, Article.citation and an accurate dateModified. All four describe something a good page already contains, and the first is a supported Google feature almost no marketing site emits.

Dataset is the interesting one. It is a supported Google feature through Dataset Search, it takes a name, a description, a temporal coverage, a creator and a licence, and it is the honest description of any page that publishes a measured aggregate. Marketing sites almost never emit it, because marketing sites almost never publish measured aggregates.

Oppira emits a Dataset node on the pages that publish our own aggregate competitor measurements, with the sample size and the as-of date stated in the visible text next to every figure. To be clear about what that buys: it makes the page eligible for Dataset Search and it is semantically true. There is no published evidence that Dataset markup increases AI citations, and anyone claiming otherwise is guessing.

The other three are ordinary and still commonly missing:

  • Person on the author, with jobTitle, url and sameAs. Cheap, and the structured half of a named accountable byline.
  • Article.citation as a CreativeWork per external source, mirroring the sources list rendered on the page.
  • An accurate dateModified that matches the visible last-updated date and the sitemap lastmod. Freshness was a unanimous gatekeeper in the SIGIR 2026 trials, so a markup date that disagrees with the page is worse than no markup.

6. Does the markup have to match what is visible on the page?

Yes. Google states that all the content in your markup must also be visible on the web page, which rules out emitting an answer, a date or an author in JSON-LD that a reader cannot find in the rendered text.

This rule quietly invalidates a whole family of clever tactics. Stuffing extra questions into FAQPage markup that do not appear on the page, emitting a fresher dateModified than the visible date, or listing an author who has no byline are all violations of it, and all three are common.

There is a second reason to follow it beyond compliance. Engines quote rendered text, and the Semrush factors that associated positively with citation, clarity and summarisation at plus 32.8% and visible question-and-answer format at plus 25.5%, are properties of the text a human sees. Markup describing content that does not exist cannot be quoted, because there is nothing there to quote.

So treat JSON-LD as a mirror. Anything in it should be a machine-readable restatement of something already on the page, generated from the same source, which also means it cannot drift out of sync when the page is edited.

7. How do I check my structured data is present and valid?

Fetch the page without JavaScript and confirm the JSON-LD is in the raw HTML, then validate it, then check the Search Console reports for the types that still have them.

Four checks, in this order. The first one catches the failure that makes the other three irrelevant.

  1. Fetch the raw HTML and search it for the JSON-LD. Request the page as a crawler would, with no JavaScript execution, and confirm the script block is in the response body. Markup injected client-side is invisible to every AI crawler, because none of them run JavaScript.
  2. Validate the syntax and the required properties. Run the markup through the Schema Markup Validator and through the Rich Results Test for the types Google still supports. Expect FAQPage and HowTo to be unrecognised, which is the correct current behaviour rather than a bug in your markup.
  3. Diff the markup against the visible page. Compare every field against the rendered text: headline, dateModified, author name, image, and every question in any Q&A markup. Anything present in JSON-LD and absent from the page violates Google guidance and should be deleted or rendered.Generating the markup from the same data object that renders the page removes this whole class of drift.
  4. Watch the reports that still exist. Check the Search Console enhancement reports for Breadcrumbs, Articles and Datasets. FAQ reporting was dropped in June 2026 and the Search Console API stops accepting it in August 2026, so its absence is expected.

What you will not find is a report telling you whether markup helped an AI answer. No engine exposes that, so resist the temptation to infer one from a citation that arrived the same week you shipped some JSON-LD.

Key Takeaways

No schema is required for AI answers

Google states structured data is not required for generative AI search and that there is no special schema.org markup to add for its AI features.

FAQPage and HowTo are finished as levers

FAQ rich results stopped appearing on 7 May 2026 and HowTo was deprecated in September 2023. Neither is in the structured data search gallery.

The schema-to-citation correlation is confounded

Citera found schema on 69 to 72% of articles at every ranking position with no gradient, while the SIGIR 2026 controlled test found formatting negligible.

Dataset is the underused supported type

Any page publishing measured aggregates can honestly emit Dataset, which feeds Google Dataset Search. Almost no marketing site does.

Markup must mirror the visible page

Google requires all content in your markup to be visible on the page, which rules out hidden questions, inflated dates and phantom authors.

Client-injected JSON-LD does not exist

No major AI crawler executes JavaScript, so markup added after page load is invisible to all of them regardless of how valid it is.

Frequently Asked Questions

Sources

  1. AI features and your website Google, July 2026.Primary. The indexing and snippet eligibility rule, with no markup requirement attached.
  2. Optimizing your website for generative AI features on Google Search Google, July 10, 2026.Primary. States that structured data is not required for generative AI search and that no special schema.org markup is needed.
  3. Structured data search gallery Google, July 2026.Primary. The authoritative list of supported types. Note the absence of FAQPage, HowTo and ClaimReview.
  4. Changes to HowTo and FAQ rich results Google, August 2023.Primary. The HowTo deprecation notice and the first stage of the FAQ rollback.
  5. Google drops FAQ rich results from Search Search Engine Journal, May 2026.Trade press reporting the May 2026 documentation change, with the June and August 2026 deprecation dates.
  6. Top ways to ensure your content performs well in Google's AI experiences Google, May 2025.Primary. Source of the rule that all content in your markup must also be visible on the page.
  7. How We Built a Content Optimization Tool for AI Search Semrush, August 2025.11,882 prompts and 304,805 cited URLs. Observational. Source of the structured data, clarity, E-E-A-T and Q&A association figures.
  8. An Analysis of 350,000 B2B SaaS Articles Citera, May 2026.Vendor research with disclosed limitations. Source of the finding that schema is present at 69 to 72% of articles at every ranking position with no gradient.
  9. What Gets Cited: Competitive GEO in AI Answer Engines ACM SIGIR 2026 (Vishwakarma, Kumar, Jamidar), May 26, 2026.Controlled paired test across 252,000 trials and six models. Found formatting and structure negligible and recency a unanimous gatekeeper.
  10. AI Answer Engine Citation Behavior: GEO16 Framework arXiv (Kumar, Palkhouski), September 2025.1,702 citations across 16 B2B SaaS verticals. Observational and single-point-in-time, as the authors state. Source of the semantic HTML and structured data correlations.
  11. GEO: Generative Engine Optimization ACM SIGKDD 2024 (Aggarwal et al.), June 2024.Controlled experiment over roughly 10,000 queries. Source of the Cite Sources result. The generative engine was a research harness, so read it as evidence about model preference rather than production retrieval.
  12. We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read Ahrefs, May 2026.Log-based analysis of 137,210 domains. The definitive dataset on whether llms.txt is fetched at all.
Oppira

Turn reading into a reaction

Oppira watches your competitors, keeps your playbook current, and drafts the response. Start free and see your market clearly by tomorrow.