What Is Answer Engine Optimisation, and How Does It Differ From SEO?
The Parts With Evidence Behind Them, and the Parts That Are Folklore
Short answer
Answer engine optimisation (AEO) is the practice of making a page retrievable and quotable by AI answer engines. It shares the technical prerequisites of SEO, indexing and crawlability, but the unit of optimisation is a passage rather than a page, and recency and concrete specifics act as gatekeepers rather than tiebreakers.
1. What is answer engine optimisation?
Answer engine optimisation is the work of getting a page retrieved and quoted by AI answer engines such as ChatGPT Search, Google AI Mode, Perplexity and Claude web search. The output is a citation, not a ranking position.
- Answer engine optimisation
- The practice of making a page retrievable by AI answer engines and writing its claims so that a single passage can be lifted out and attributed.
- Also known as: AEO, Generative engine optimisation, GEO
The naming is unsettled and worth knowing about. Practitioners mostly say AEO, while the academic literature says GEO, from Aggarwal and colleagues, "GEO: Generative Engine Optimization", published at ACM SIGKDD 2024. Both names describe the same goal, so treat them as synonyms and ignore anyone selling a distinction between them.
A citation is a concrete, observable thing. When an assistant answers a question, it lists the pages it used, and in some cases the exact span of text it took. Being cited means your URL appears in that list. Being described without a citation is a different outcome, and it comes mostly from what other sites say about you rather than from your own pages.
2. How is answer engine optimisation different from SEO?
SEO optimises a page to rank against other pages. Answer engine optimisation optimises a passage to be quoted. The crawling and indexing prerequisites are identical, but length and keyword density stop being levers.
| Dimension | Classic SEO | Answer engine optimisation |
|---|---|---|
| Unit optimised | The page, ranked against other pages | The passage. Anthropic's web search tool returns up to 150 characters of cited text per result |
| Success metric | Position in the result list, then clicks | Whether the page is cited at all, and in which citation slot |
| Technical prerequisite | Indexed and crawlable | Indexed and crawlable, plus present in raw HTML before any JavaScript runs |
| Word count | Longer pages tend to rank across pages 1 and 2 | Spearman 0.04 against citation position across 174,048 cited pages (Ahrefs) |
| Recency | A tiebreaker on most queries | A gatekeeper. Claude web search hands the model a page age with every single result |
| Relationship to rankings | Rankings are the goal | 12% of AI-cited URLs rank in Google top 10 for the original prompt (Ahrefs, 15,000 long-tail queries) |
| Markup | Rich results depend on specific schema types | Google states no special schema.org markup is needed for its generative AI features |
The practical consequence is a change of unit, not a change of craft. A comprehensive page that buries its answer in paragraph six ranks fine and gets quoted rarely, because there is no sentence on it that survives extraction. A short page with one exact, dated, self-contained claim does the opposite.
3. Which engines does answer engine optimisation target?
Six surfaces matter in 2026: Google AI Overviews, Google AI Mode, ChatGPT Search, Perplexity, Claude web search and Microsoft Copilot. Each retrieves from a different index, so visibility in one guarantees nothing about the others.
| Engine | Retrieval path | Crawler to allow | Overlap with Google top 10 |
|---|---|---|---|
| Google AI Overviews | Google's Search index, plus query fan-out | Googlebot | 38% of citations rank in the top 10 (Ahrefs, March 2026) |
| Google AI Mode | Same index, wider fan-out, custom Gemini model | Googlebot | 54% domain overlap (Semrush) |
| ChatGPT Search | OpenAI's own search index | OAI-SearchBot | Roughly 7 to 8% for the original prompt (Ahrefs) |
| Perplexity | Own index plus live fetch | PerplexityBot | 28.6%, the highest of any engine measured (Ahrefs) |
| Claude web search | Search executed server-side by Anthropic | Claude-SearchBot | Not measured in any dataset published so far |
| Microsoft Copilot | Grounded on Bing search results | Bingbot | Roughly 8.6% (Ahrefs) |
Two of these are gates rather than opportunities. OpenAI documents that sites opted out of OAI-SearchBot "will not be shown in ChatGPT search answers", and Microsoft documents that Copilot Search is grounded on Bing results. So a robots.txt line and a Bing index entry decide two of the six surfaces before any writing happens.
Google is the only engine where the eligibility bar is published in full. A page has to be indexed and eligible to be shown in Google Search with a snippet, and Google states there are no additional requirements and no special optimisations necessary.
4. Is answer engine optimisation a separate discipline or a rebrand of SEO?
Mostly a rebrand, with three genuine additions: raw-HTML retrievability, passage-level writing, and honest freshness. The best controlled benchmark to date found traditional SEO more effective than AI-specific tactics.
The strongest evidence against AEO as a separate craft comes from a benchmark built to test it. Puerto and colleagues, "C-SEO Bench: Does Conversational SEO Work?", published at NeurIPS Datasets and Benchmarks 2025, concluded that "most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking", while traditional SEO was "significantly more effective". The same paper showed that whatever gains exist decay as adoption rises.
Three things do genuinely change, and all three are cheap:
- Raw-HTML retrievability. Vercel and MERJ, analysing roughly 1.3 billion AI crawler requests in December 2024, found that none of the major AI crawlers execute JavaScript. Client-rendered text is invisible to them even when Googlebot handles it fine.
- Passage-level writing. Anthropic documents that its web search tool returns up to 150 characters of cited text per citation, which is roughly one sentence. A claim that only makes sense across three sentences cannot be quoted whole.
- Honest freshness. Anthropic also returns a page age field with every result, so the model sees how old your page is before choosing what to quote.
5. What does the evidence say actually influences AI citations?
Four factors were unanimous gatekeepers across six models in the largest controlled test published: topic relevance, explicit concrete specifics, recency of the timestamp, and position in the retrieved list. Formatting changes were negligible.
The test is Vishwakarma, Kumar and Jamidar, "What Gets Cited: Competitive GEO in AI Answer Engines", ACM SIGIR 2026. It ran 252,000 trials across six large language models and 18 content factors, injecting paired documents into retrieval context with brands anonymised and source order counterbalanced. The four gatekeepers above had odds ratios far above 100.
Secondary factors were significant in four or more of the six models, in descending confidence:
- Query-term presence in the text, with odds ratios from 5.99 to 40.0.
- Completeness of specifications, and inclusion of explicit comparative analysis.
- Confident rather than hedged language, with odds ratios from 2.67 to 754.
- Supporting evidence and internal consistency.
- Depth of coverage, and strength of social proof.
What the same test found negligible is the interesting half: formatting and structure changes, paragraph organisation, and neutral versus promotional tone. That inverts most published AEO advice, which is largely about reformatting.
The honest caveat is the domain. The trials sat in a product-review setting with documents injected into context rather than retrieved live, so the ranking of factors is better evidence than the absolute effect sizes.
6. Which answer engine optimisation claims should you ignore?
Ignore llms.txt, content chunking, FAQ schema as a lever, and every specific word-count threshold. Google states directly that Search does not use AI text files and that content need not be broken into small pieces.
| Circulating claim | Status | What the traceable evidence says |
|---|---|---|
| Pages over 2,900 words are 59% more likely to be cited by ChatGPT | Unverified | No sample size, method or date has ever been published. Ahrefs measured Spearman 0.04 between word count and citation position across 174,048 cited pages |
| AI Overviews extract passages of 134 to 167 words | Unverified | No traceable primary source. Semrush, across 304,805 cited URLs, found passage length negligible |
| 44.2% of citations come from the first 30% of the page | Unverified | No primary publication exists. Front-loading answers is well supported by other evidence, but this specific number is not citable |
| FAQ schema is weighted about 40% higher in ChatGPT source selection | Unverified | OpenAI publishes nothing about source weighting, and Google stopped showing FAQ rich results on 7 May 2026 |
| Schema-marked pages are cited 2.3 times more often | Unverified | Citera found schema markup present on 69 to 72% of articles at every ranking position, with no gradient at all |
| An llms.txt file improves AI retrieval | Wrong | Google states Search does not use AI text files. Ahrefs checked 137,210 domains in May 2026 and 97% of published llms.txt files received zero requests |
7. Where should a small team start with answer engine optimisation?
Start with retrievability, then evidence, then freshness. Fetch a page with JavaScript disabled, allow the search crawlers, add one specific nobody else has published, and put an honest last-updated date on it.
Six steps, in order. The first two take an afternoon and gate everything after them.
- Fetch your own page with JavaScript disabled. Request the raw HTML a crawler receives and read it. The sentence you want quoted has to be in that response, because no major AI crawler executes JavaScript. An empty container plus a script is an invisible page.
- Allow the search crawlers in robots.txt. Confirm Googlebot, Bingbot, OAI-SearchBot, PerplexityBot and Claude-SearchBot are all permitted. OpenAI states that sites opted out of OAI-SearchBot will not appear in ChatGPT search answers, which is a self-inflicted exclusion from one of the six surfaces.
- Write one self-contained answer per heading. Make every heading a question a person would ask, then answer it in the first sentence underneath with no pronoun pointing backwards. Read the sentence with the heading removed. If it stops making sense, it is not quotable yet.
- Add a specific only you can publish. Insert at least one checkable thing: a figure with its method and date, a limit of a public data source stated exactly, or two named options compared with real values. Concrete specifics were a unanimous gatekeeper in the SIGIR 2026 trials.A measured number with a stated sample size has no substitute source, which is precisely when attribution happens instead of paraphrase.
- Date the page and record what changed. Show a visible last-updated date, emit a matching dateModified, and keep a one-line changelog entry per revision. Republishing an unchanged page under a new date is a claim the text itself contradicts.
- Sample your visibility repeatedly, never once. Ask each target question 10 to 30 times per engine before believing a result. Schulte and colleagues show single measurements of AI visibility are unreliable because model output is stochastic, so a one-run report is noise.
What deliberately is not on that list: publishing an llms.txt file, splitting pages into chunks, adding FAQ markup, and expanding pages to hit a word count. All four are common first moves and none has measurable support.
Recommendation questions are the exception to everything above, because they are decided off your site. Oppira tracks which brands and domains get named across the public sources those answers lean on, which turns "why do assistants recommend someone else" into a list you can act on rather than a guess.
Key Takeaways
The unit is the passage, not the page
Anthropic's web search tool returns up to 150 characters of cited text per result, so a quotable claim is roughly one self-contained sentence.
Four factors were unanimous gatekeepers
Across 252,000 trials and six models at SIGIR 2026: topic relevance, concrete specifics, recency, and position in the retrieved list.
Formatting is not the lever
The same controlled test found formatting, structure and paragraph organisation negligible, which inverts most published AEO advice.
Word count is not the lever either
Ahrefs measured Spearman 0.04 between word count and citation position across 174,048 cited pages. No threshold effect has ever been demonstrated.
Retrievability comes before writing
No major AI crawler executes JavaScript, and opting out of OAI-SearchBot removes you from ChatGPT search answers entirely.
Traditional SEO beat AI-specific tricks in the benchmark
C-SEO Bench at NeurIPS 2025 found conversational SEO methods largely ineffective and often harmful to ranking, with traditional SEO significantly more effective.
Frequently Asked Questions
Sources
- AI features and your website Google, July 2026.Primary. States the indexing and snippet eligibility bar for AI Overviews and AI Mode, and documents query fan-out.
- Optimizing your website for generative AI features on Google Search Google, July 10, 2026.Primary. Directly rejects llms.txt, content chunking, AI-specific writing styles and special schema.
- What Gets Cited: Competitive GEO in AI Answer Engines ACM SIGIR 2026 (Vishwakarma, Kumar, Jamidar), May 26, 2026.252,000 trials, six models, 18 factors, paired injection with counterbalanced order. Highest causal quality available, but set in a product-review domain.
- C-SEO Bench: Does Conversational SEO Work? NeurIPS Datasets and Benchmarks 2025 (Puerto, Gubri, Green, Oh, Yun), June 2025.Finds conversational SEO methods largely ineffective and often harmful to ranking, with traditional SEO significantly more effective.
- Short vs. Long Content in AI Overviews Ahrefs, January 2026.560,346 AI Overviews and 174,048 cited pages with word counts. Source of the Spearman 0.04 word-count result.
- Only 12% of AI Cited URLs Rank in Google's Top 10 Ahrefs, January 2026.15,000 long-tail queries across ChatGPT, Gemini, Copilot and Perplexity. Source of the per-engine overlap figures.
- We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read Ahrefs, May 2026.Log-based rather than survey-based, which is why it is the definitive llms.txt dataset.
- How We Built a Content Optimization Tool for AI Search Semrush, August 2025.11,882 prompts and 304,805 cited URLs. Observational, and the clearest evidence that passage length and word count are negligible.
- An Analysis of 350,000 B2B SaaS Articles Citera, May 2026.Vendor research for their own platform, with limitations disclosed. Source of the finding that schema markup shows no gradient by ranking position.
- Web search tool Anthropic, July 2026.Primary. Documents the page age field on every result and the 150 character cited text cap on every citation.
- The rise of the AI crawler Vercel with MERJ, December 17, 2024.Roughly 1.3 billion AI crawler requests. The no-JavaScript-rendering finding. Dated December 2024, so worth re-testing rather than assuming.
- Don't Measure Once: Measuring Visibility in AI Search arXiv (Schulte, Bleeker, Kaufmann), January 2026.Shows single measurements of AI visibility are unreliable and recommends roughly 10 to 30 samples per prompt with confidence intervals.
Explore More
Related analyses, benchmarks, and industry insights
Related Guides
Glossary Terms
Turn reading into a reaction
Oppira watches your competitors, keeps your playbook current, and drafts the response. Start free and see your market clearly by tomorrow.