Generative engine optimization: what the controlled evidence actually shows in 2026
One 252,000-trial experiment, a 45-study survey and four crawler-log and correlation studies, read from the source tables. The factors that decide a citation are not the ones being sold, and the one most often sold measured as no effect.
Every agency deck on generative engine optimization has the same slide: add headings, add schema, add an llms.txt file, write shorter, sound confident. The pitch is that AI answer engines reward tidy pages. We read the primary studies behind that pitch, from the source tables rather than the summaries, and the slide does not survive.
The largest controlled experiment on record, 252,000 paired trials across six language models and 18 content factors, found that organised sections against dense prose landed at an odds ratio of 0.78 to 1.68 depending on the model, a range that straddles no effect [source]. The authors’ own words: “Formatting changes showed minimal return and can be deprioritized” [source].
What did decide the citation was not on the slide at all.
What generative engine optimization means
Generative engine optimization is the practice of making a page more likely to be retrieved and cited by an AI answer engine such as ChatGPT, Perplexity, Gemini or Google’s AI Overviews. It is not one ranking task. A 2026 survey of 45 studies describes it as “a stochastic, partially observable pipeline spanning search activation, crawling and indexing, retrieval, reranking and context allocation, citation, prominence, factual absorption, fidelity, and user behavior” [source]. A page can win the last stage and lose the first three.
The four gatekeepers, and the one factor that did nothing
The experiment worth reading is Vishwakarma, Kumar and Jamidar’s two-document testbed, presented at SIGIR 2026. Two sources are injected into a retrieval-augmented answer, differing in exactly one attribute, counterbalanced for order and with brands anonymised, and the measurement is which one the model cites first. Six models, 18 factors, 252,000 trials [source].
Four factors were unanimous across all six models, with odds ratios the authors cap at “>>10k” because the losing variant almost never won:
| Factor | Odds ratio across six models | The authors’ category |
|---|---|---|
| Topic match | 221 to >10k | Gatekeeper |
| Position in the model’s context | 1,795 to >10k | Gatekeeper |
| Recent timestamp against a 2019 one | 14.4 to >10k | Gatekeeper |
| Price stated, on product queries | 6.26 to >10k | Gatekeeper |
| Specifications present | 8.63 to >10k | Differentiator |
| Comprehensive coverage | 3.98 to >10k | Differentiator |
| Claims carry evidence | 2.09 to >10k | Differentiator |
| Query terms appear on the page | 5.99 to 40 | Differentiator |
| Unhedged language | 2.67 to 754 | Differentiator |
| Comparisons included | 1.61 to 7.45 | Differentiator |
| Internally consistent | 1.74 to 4.09 | Differentiator |
| Organised sections against dense prose | 0.78 to 1.68 | No consistent effect |
All values from Table 2 of the paper [source]. The “>>10k” figures are what the authors call quasi-separation: “a decisive win for variant A rather than a finely resolved numeric ratio” [source].
Read the bottom row. Then read the top four. The factors that decide a citation are whether the page is about the thing asked, where the retriever placed it, whether it carries a live date, and whether it names a price. None of those is a formatting choice. On the factor every GEO vendor sells first, the paper’s finding is that “LLMs parse content regardless of visual organization” [source].
Social proof, the other staple of the deck, reached significance in two of the six models [source]. Whether a date beats no date at all was weak: 1.15 to 13.0. A missing date and a stale date were near-equivalent at 1.31 to 2.32 [source]. The lesson is not “add a date”. It is “a page that reads as 2019 loses to one that reads as this year”.
Structure: two studies, two answers, and why both are right
We published in September that structure decides whether a passage is quoted, resting in part on the GEO-SFE preprint, which reports a 17.3% improvement in citation rate across six generative engines when content is re-engineered structurally [source]. That finding and the one above look like a contradiction. They are not, and the reconciliation matters more than either number.
GEO-SFE varied structure on real pages retrieved for real queries: whether a passage is self-contained enough to be lifted from a page and placed in front of the model. Vishwakarma’s testbed hands the model two documents that are already in front of it and asks which it cites. The first study is about getting a passage into the context. The second is about what happens once two passages are there. Structure helps you arrive. Once you have arrived, topic, position, date and price decide, and the tidiness of your headings does not.
We show the two results side by side rather than picking one. Anyone selling structure as a citation lever without saying which stage they mean is selling the wrong half of the pipeline.
Visibility is a product of four probabilities
The survey’s central point is that every stage multiplies. Its evidence, per stage:
Stage 1, does an AI answer appear at all. Across 55,393 trending queries over 40 days, an AI Overview appeared on 13.7% of them, rising to 64.7% when the query was phrased as a question [source].
Stage 2, is the page retrieved. In one 2026 study of what AI systems tried to fetch, 27.1% of URLs “were not scraped because they were inaccessible, removed, or non-textual” [source]. And the crawler logs are unambiguous on rendering. Vercel and MERJ, across 569 million GPTBot fetches and 370 million from Anthropic’s crawler in a single month: “None of the major AI crawlers currently render JavaScript. This includes: OpenAI (OAI-SearchBot, ChatGPT-User, GPTBot), Anthropic (ClaudeBot), Meta (Meta-ExternalAgent), ByteDance (Bytespider), Perplexity (PerplexityBot)” [source]. Two do render: Google’s Gemini through Googlebot, and AppleBot through a browser-based crawler [source]. GPTBot fetched JavaScript files on 11.50% of requests and Claude on 23.84%; neither executed them [source]. A fact your page only states after JavaScript runs does not exist to most of the machines deciding whether to cite you.
Stage 3, is it cited once retrieved. The gatekeepers above. This is the only stage most GEO advice addresses.
Stage 4, does anyone click. The survey grades the evidence “very low”: one suggestive quasi-experiment and industry claims, causality not established [source]. The one conversion study that ran a placebo check returned p = 0.16 [source].
The trap sits between stages two and three. In an arena study across 171,003 documents and 2,700 queries, optimising the body of a page for citation “reduces average top-20 presence by approximately 9%, top-10 presence after reranking by 16%, and final citation by 6%” [source]. Rewriting to win the visible stage lost the invisible one, and the net was negative.
Being known is not being recommended
The standard executive test is to open ChatGPT and ask what it knows about the company. It returns a flattering paragraph. That paragraph measures recognition, and recognition is near-universal.
A January 2026 preprint tested 112 startups from Product Hunt’s top 500 across 2,240 queries. Asked about a product by name, ChatGPT recognised it 99.4% of the time and Perplexity 94.3%. Asked the question a buyer would actually ask, without the name, the products appeared 3.32% and 8.29% of the time [source]. The same study scored each startup’s site on GEO practices and found the score “showed no correlation with actual discovery rates”; what did predict discovery on Perplexity was referring domains and Product Hunt ranking [source].
One author, one preprint, so this is a structural point rather than a settled number. The structure is the useful part: the number that pays is the second one, and almost nobody measures it.
The strongest signals are not on the page
Ahrefs’ correlation studies across 75,000 brands say the same thing from the other side. Against AI Overview visibility, branded web mentions correlated at a Spearman 0.664 and backlinks at 0.218 [source]. Extended to ChatGPT and AI Mode as well, YouTube brand mentions were the single strongest signal at approximately 0.737, branded web mentions ran 0.656 to 0.709 by platform, and backlinks 0.194 to 0.268 [source]. Both reports carry the authors’ own warning that correlation is not causation [source].
Content length, the metric most content programmes optimise, sat at a Spearman 0.04 against citation position across 174,048 pages cited in 560,346 AI Overviews. More than half of citations, 53.4%, went to pages under 1,000 words; 16.0% went to pages over 2,000 [source].
A page-template programme raises the ceiling. It does not create the mentions. A proposal that presents on-page work as the whole answer is mispricing the problem.
Sold hard, supported weakly
| Tactic | The evidence | Verdict |
|---|---|---|
| Schema / JSON-LD for citations | 1,885 pages tracked adding schema: +2.4% on AI Mode and +2.2% on ChatGPT, both “statistically indistinguishable from zero”; a 4.6% decline on AI Overviews. Five AI systems tested during live retrieval “extracted only visible HTML content” [source] | No effect |
| llms.txt | 137,210 domains: 28% publish one, 97% of those received zero requests in May 2026 [source] | No effect |
| Longer content | Spearman 0.04 against citation position, 174,048 pages [source] | No effect |
| Writing simpler | The 2024 KDD paper’s “Easy-to-Understand” rewrite scored 22.0 against a 19.3 baseline on position-adjusted word count, a 14.0% gain, against 42.6% for adding quotations and 32.8% for adding statistics [source] | Marginal |
| Authoritative tone | The survey grades it “weak and unstable” and warns: “do not conflate confidence with evidence” [source] | Marginal |
| Keyword stuffing | 17.7 against a 19.3 baseline, the only method in the table below baseline, an 8.3% loss; the authors: “little to no improvement” [source] | Harmful |
| Server-rendering every fact | No JavaScript execution observed across hundreds of millions of fetches by OpenAI, Anthropic, Meta, ByteDance and Perplexity crawlers [source] | Decisive |
| A real price on the page | Gatekeeper on product queries, unanimous across six models, odds ratio 6.26 to >10k [source] | Decisive |
| Verifiable dated evidence | Graded “moderate to strong” by the survey, “dependent on intent, engine, and factuality” [source]; odds ratio 2.09 to >10k in the controlled test [source] | Supported |
Two caveats the table cannot hold. The schema pages were already heavily cited before schema was added, so the study says schema does not push a cited page higher; it does not test an uncited one [source]. And the 14.0% and 8.3% figures are our arithmetic on the paper’s Table 1, not numbers the paper prints as percentages.
Where SEO and GEO disagree
One global rule would be cheaper to maintain and wrong on half the pages. The ruling changes by page type because a different mechanism decides each.
Comparison and alternatives pages go GEO first: they are the format answer engines lean on and the structure serves the organic result too. Glossary, FAQ and how-to pages go GEO first for the same reason: passage retrieval fits them and SEO gives up nothing. Product and pricing pages split: the facts and the price belong to the machine, the space above the fold belongs to conversion. Location and sector pages stay SEO first. In Whitespark’s 540-query local study, AI Overviews appeared on 68% of local business queries, and in its Houston plumber sample 60% of the citations went to third-party publishers and 40% to individual businesses [source]. That is one city and one trade, and it is the only measured split we could find. The local pack still decides the local page.
Most reported GEO results are noise
Ask the same engine the same question twice and it cites a different set of pages.
| Measurement | Value | What it forces |
|---|---|---|
| Agreement between identical daily runs, four engines, 45 days | Jaccard 0.34 to 0.42 [source] | Seven to eight repetitions per prompt as a starting point [source] |
| Decisions that change on repeated runs at temperature zero | 9% to 28% [source] | Keep the nulls in the denominator |
| Sentences fully supported by their own citation, four engines, 2023 | 51.5% [source] | Citation is not endorsement |
| URL overlap between Google organic, AI Overviews and Gemini, 11,500 queries | Jaccard 0.11 to 0.18 [source] | Every metric names its engine, surface and period |
| Generic GEO methods that transferred across domains | 3 of 54 method-domain combinations significantly positive [source] | Nothing generalises without local testing |
The survey’s conclusion is the sentence to keep: “no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior” [source]. A single before-and-after cannot separate an effect from run-to-run variance. A vendor reporting a round multiplier from one run is reporting the weather.
The proof
TIN’s own case files hold two records that bear on this. When Google launched AI Overviews to U.S. users in May 2024, the feature advised non-toxic glue on pizza and a rock a day within days, and Google’s head of Search announced more than a dozen technical improvements and new restrictions; the stage-one number above, 13.7% of queries, is what remained after that retreat case file. And the FTC’s final order against Workado records a vendor advertising 98.3% accuracy for an AI tool whose own published tests showed 53.2% on real-world text case file. A vendor’s number about what its tool does is a claim until someone opens the table. That is the whole method of this piece.
The bottom line
Get the page fetched: server-render every fact, because most AI crawlers do not run your JavaScript. Get it retrieved: match the topic in the words the query uses, and carry a date that reads as this year. Once it is in front of the model, state the price, show the specifications, put evidence under every claim, and stop hedging. Earn mentions off the page, because that is where the strongest correlations live. Measure with seven or eight repetitions and name the engine every time.
Do not pay for headings, schema, llms.txt or word count as citation levers.
Each has now been measured, and each measured as nothing.
Method
Every figure in this piece was read from the source table or the quoted sentence of the primary document, not from a summary of it. Where a figure is our arithmetic on a published table, the text says so. Where two studies disagree, both are shown. Four figures in our own first draft were wrong and were corrected before publication: the crawler-rendering claim omitted AppleBot, social proof was significant in two models rather than “two or three”, a widely repeated llms.txt fetch count had no primary source and was replaced with Ahrefs’ measured figure, and the content-length correlation is against citation position, not citation.
Sources
- 01Rahul Vishwakarma, Shushant Kumar, Ratnesh Jamidar, “What Gets Cited: Competitive GEO in AI Answer Engines”, Proceedings of the 49th International ACM SIGIR Conference, July 2026; arXiv 2605.25517, submitted 25 May 2026. https://arxiv.org/abs/2605.25517
- 02Olivier Martinez, “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)”, arXiv 2607.14035, submitted 15 July 2026. https://arxiv.org/abs/2607.14035
- 03Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande, “GEO: Generative Engine Optimization”, KDD 2024; arXiv 2311.09735. https://arxiv.org/abs/2311.09735
- 04Amit Prakash Sharma, “The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries”, arXiv 2601.00912, submitted 1 January 2026. https://arxiv.org/abs/2601.00912
- 05Junwei Yu, Mufeng Yang, Yepeng Ding, Hiroyuki Sato, “Structural Feature Engineering for Generative Engine Optimization: How Content Structure Shapes Citation Behavior”, arXiv 2603.29979, submitted 31 March 2026. https://arxiv.org/abs/2603.29979
- 06Giacomo Zecchini, Alice Alexandra Moore, Malte Ubl, Ryan Siddle, “The rise of the AI crawler”, Vercel with MERJ, 17 December 2024. https://vercel.com/blog/the-rise-of-the-ai-crawler
- 07Louise Linehan, “We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved.”, Ahrefs, 11 May 2026. https://ahrefs.com/blog/schema-ai-citations/
- 08Louise Linehan and Xibeijia Guan, “We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read”, Ahrefs, 15 June 2026. https://ahrefs.com/blog/llmstxt-study/
- 09Louise Linehan, “An Analysis of AI Overview Brand Visibility Factors (75K Brands Studied)”, Ahrefs, 26 May 2025. https://ahrefs.com/blog/ai-overview-brand-correlation/
- 10Louise Linehan, “Top Brand Visibility Factors in ChatGPT, AI Mode, and AI Overviews (75k Brands Studied)”, Ahrefs, 12 December 2025. https://ahrefs.com/blog/ai-brand-visibility-correlations
- 11Despina Gavoyannis, “Short vs. Long Content in AI Overviews: The Data Says Both Work”, Ahrefs, 3 December 2025. https://ahrefs.com/blog/short-vs-long-content-in-ai-overviews/
- 12Miriam Ellis, “New Research: AI Overviews for Local Business Searches”, Whitespark, 12 May 2025. https://whitespark.ca/blog/case-study-the-prevalence-of-ai-overviews-in-local-search/
Questions
Does schema markup help AI citations?
No measurable lift in the one controlled test on record. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 and found +2.4% on Google AI Mode and +2.2% on ChatGPT, both statistically indistinguishable from zero, and a small decline of 4.6% on AI Overviews. Five AI systems tested during live retrieval consumed none of the markup.
Do AI crawlers run JavaScript?
The crawlers from OpenAI, Anthropic, Meta, ByteDance and Perplexity do not, across hundreds of millions of logged fetches. Google's Gemini uses Googlebot and renders JavaScript, and AppleBot renders it too. Anything a page only says after JavaScript runs is invisible to most AI fetchers.
Does llms.txt work?
Not on the evidence. Across 137,210 domains, 28% publish an llms.txt file and 97% of those files received zero requests in May 2026. Adoption is real; consumption is not.
What actually decides whether an AI engine cites a page?
In a 252,000-trial controlled experiment across six models, four factors were unanimous: the page matches the topic, it sits early in the model's context, it carries a recent date rather than a stale one, and on commercial queries it states a price. Structure versus dense prose measured at an odds ratio of 0.78 to 1.68, straddling no effect.
This is analysis, not a verified outcome. It carries no verification badge and never will. The proof lives in the case files, where every figure is checked against the public record and the method is printed on the page.