AI search optimization: what ChatGPT's retrieval stream shows, and the numbers nobody can source
A viral analysis of ChatGPT's server-sent events says the engine ranks passages, not pages. We checked it line by line. The mechanism holds up across four independent observers. Two of the numbers being quoted from it are not in it.
An analysis of ChatGPT’s retrieval system has been circulating for the past week under the word “leak.” It reports that for a single prompt, ChatGPT ran 5 search rounds, wrote 18 hidden queries, made 50 engine calls, pulled 228 results, fetched 223 URLs, and cited 16 [source].
The mechanism it describes is real, corroborated by three other researchers working independently, and it should change how you write.
Two of the numbers now being quoted from it are not in it. We went looking for them and they are not there.
What AI search optimization actually means
AI search optimization is the practice of making individual passages retrievable and quotable, rather than making a page rank. An answer engine fetches your page, splits it into blocks, scores each block on its own, and shows the model a handful of them. The rest of the page is never read by anything that decides what to say. Your page is not the unit of competition. Your paragraphs are, one at a time, and most of them lose.
That is the claim. Here is what actually supports it.
There was no leak, and that is the good news
What was analysed is the server-sent events stream, the channel ChatGPT uses to push its answer to your browser while it types. Open developer tools, ask a question, watch the network tab. The author describes pulling this data for several days and hitting rate limits [source].
Calling it a leak inverts its evidentiary status. A leak is a secret you take on faith. Client-side telemetry is a reproducible observation: anyone can run it again, and anyone can contradict it. That is the stronger position, and it is the one the word “leak” gives away.
The mechanism: four observers, one conclusion
The passage-level finding does not rest on one person.
First, who is observing. Yeşilyurt is a GEO researcher at Peec AI, an AI visibility tracker that announced in June 2026 it had assembled a team to decode how answer engines recommend brands [source]. The first post discloses it: the prompt under the microscope was “best AI visibility tools”, the fetched results included the team’s own Peec AI product page, and the collaborators named are Peec colleagues [source]. That makes the analysis vendor research, run on a prompt about the vendor’s own category. It does not make it wrong, and the disclosure is the author’s own. It is the reader’s to weigh, and it belongs at the top of any piece that quotes the work as often as this one does.
Yeşilyurt reports the page render cut into blocks of roughly 170 words, each scored, with the model seeing one to three blocks per source, a should_fetch decision on each result, and a 150,000-line JSON dump from one prompt [source]. The second post names the fields: snippet::parts holding three passages from a page and snippet::scores holding one neural relevance score for each [source].
Filippo Danesi, writing in July 2026, relays Suganthan Mohanadasan’s network-traffic analysis finding that citations bind to a specific sentence rather than to the page, alongside AirOps’ finding that 85% of retrieved pages never appear in the final answer [source]. Danesi’s piece also draws on Yeşilyurt’s earlier work for the size of the retrieval window and the rank-fusion method, so the independence that counts here is narrower than the byline: the sentence-level finding is Mohanadasan’s own traffic analysis, not a restatement [source].
Olivier de Segonzac of RESONEO, reporting in August 2026, describes OpenAI’s own retrieval hub holding a cache of full pages converted from HTML to Markdown, and counts 61,332 URLs surfaced into the sources sidebar against 759 pages actually opened [source].
And a controlled experiment, rather than an observation, reaches the same place from the other direction. The GEO-SFE preprint held semantic content constant and varied structure alone across 200 articles, 377 real-world queries and six generative engines, 2,400 test cases in total, and measured a 17.3% improvement in citation rate at p<0.001 [source].
Four methods, one answer: structure decides whether you are quoted.
The survival rates do not reconcile, and nobody says so
Here is what gets lost when a single post goes viral. The three sets of figures above are being read as one story. They are three instruments that disagree.
| Observer | What was counted | Survival rate |
|---|---|---|
| Yeşilyurt, Sept 2026 | 223 URLs fetched, 16 cited, one prompt | 7.2% of fetched |
| AirOps, via SERPsecrets | 85% of retrieved pages never appear | 15% of retrieved |
| RESONEO, Aug 2026 | 61,332 URLs surfaced, 759 pages opened | 1.2% opened |
These are not the same quantity. “Returned by an engine” is not “fetched” is not “opened” is not “surfaced in the sidebar,” and “cited” is not “was the lead source behind a citation.” RESONEO’s own set has 5,032 URLs recorded as the lead source behind a citation against 759 pages actually opened, a pair that cannot be composed without knowing exactly what each term counts [source].
We are not averaging them. The direction is agreed and the magnitude is not, which is a normal state for evidence and an unusual one for a strategy post. If you are told that “7% of pages survive,” ask which denominator, on how many prompts.
The two numbers that are not in the source
This is the part we went looking for, and it is the reason this post exists.
Two claims reached us in secondhand summaries of the analysis, and both are being used to argue that the scores are worthless. Neither is in the source. The distortion is the summarisers’, not the author’s: if you are quoting either claim to say the research is wrong, you are quoting something it never said. We could not find either claim in any indexed public write-up, so we cannot link the distortion; we print it so that a reader who meets it knows what to check.
The first is that the top-scoring block is normalised to exactly 1.0000, which would mean a score of 1.0 says “best in this batch” rather than “excellent,” and that comparing scores across queries is meaningless. The second is that the findings were verified across 76 pre-indexed pages.
We searched the full text of both posts and the author’s own longer write-up on his ranking research [source]. Neither claim appears in any of them.
What the second post does publish is three neural relevance scores from one page, a SoundGuys headphone comparison: 0.9817 on a battery-life passage, 0.9937 on a “what to buy instead” passage, and 0.9984 on a passage covering Sennheiser and Monoprice alternatives [source].
Three values, from one page, on one query. That query is a shopping query, a comparison of noise-cancelling headphones under $300, and the first post separately notes that shopping and local queries carry their own result types [source]. The author adds his own caveat on the scores: they change with A/B testing variants, and he is still working on it [source].
So the honest reading is narrower than either side of the argument. There is no published evidence that scores are max-normalised, and there is no published sample large enough to read anything into the gaps between them. Three thousandths of a point, measured once, in the most atypical vertical, is not a thing anyone can optimise for. Keep the mechanism. Do not build a process on the scores.
Do not write 170-word blocks because a stream said so
The most quoted operational takeaway from the analysis is the 170-word block. Treat it carefully, because the second-most-detailed observer measured something an order of magnitude smaller: RESONEO puts snippets at roughly 200 characters in ChatGPT’s instant mode, drawn from the H1 and whatever visible text sits around it, and reports one case that came out 100% table of contents and 0% content [source].
Those two are probably not contradicting each other. They are probably describing different retrieval modes of the same product. But neither author reconciles them, which means anyone handing you 170 words as a spec is generalising from one mode of one engine on one day.
The range worth acting on is 150 to 300 words, and the reason is narrower than this piece first printed. GEO-SFE recommends that range as a design principle in its optimisation section, and what it tested was its structural framework as a whole: 200 articles, each in a baseline and an optimised version, across six engines, 2,400 test cases, a 17.3% lift in citation rate at p<0.001 [source]. It did not test paragraph length on its own, and no table in the paper stratifies citation rate by length. The 31% attention-degradation figure it prints is cited from Liu et al.’s 2023 study of position bias in long contexts [source], and the 23% figure for chunks under 150 words carries no citation at all [source]. So the range sits inside a tested bundle, which beats a number read off a stream, and it is not a measured optimum.
Weigh the study itself the way this piece weighs the survival rates. It is an arXiv preprint submitted on 31 March 2026, not peer-reviewed work. Its citation rate is queries citing the content over total queries, on one unreplicated run. And its six engines include Google SGE and Bing Chat, product names retired long before submission, with no date given for when the runs were made [source]. One controlled study outranks four field observations on method. It does not settle the question.
And no, passage ranking is not new
Google announced passage indexing in October 2020 and turned it on for US English queries on 10 February 2021, saying it expected the change to affect 7% of queries across all languages once fully rolled out. Google’s own description at the time: “we’ve recently made a breakthrough in ranking and are now able to not just index web pages, but individual passages from the pages” [source].
Five and a half years later the same architecture is being announced as a leak. What is genuinely new is not the ranking of passages. It is that you can now watch it happen in your own browser.
Schema markup for AI search: the study that says stop
If you are choosing between spending on structure and spending on markup, one controlled test settles it.
Ahrefs tracked 1,885 URLs that added JSON-LD schema between August 2025 and March 2026, against roughly 4,000 matched control pages, measured over 30 days before and after. ChatGPT citations moved 2.2% and Google AI Mode 2.4%, both indistinguishable from zero. AI Overviews fell 4.6%, a statistically significant decline, though both the treated and control groups were caught in a broader platform-wide drop unrelated to schema [source].
Read that carefully in both directions. It does not show that schema hurts. It shows that adding schema to a page that already had citations did not produce more of them. Structure pays. Markup is a symptom of a well-run site, not the cause of its citations.
The proof: why an unchecked number about AI travels for years
TIN’s registry exists because of what happens next to a number like the ones above.
Workado marketed an AI Content Detector as 98.3% accurate at spotting AI-generated text. The FTC’s complaint alleges the model behind it was a publicly available open-source detector built by students in Norway for an undergraduate thesis and trained only on academic abstracts, and that its developers’ own published tests showed 53.2% correct detection of AI-generated non-academic text, which the FTC described as barely better than a coin toss. The correction arrived as final consent order C-4822 on 28 August 2025, from a regulator, not from the market that had been quoting the number (case file).
The second case file is the mechanism failing in public. In May 2024, days after Google launched AI Overviews to US users, the feature advised putting non-toxic glue on pizza to keep the cheese on and stated that geologists recommend eating one rock a day. Google’s head of Search acknowledged the failures and announced more than a dozen technical improvements (case file).
Both of those are the same lesson from opposite ends. The passage is the unit of citation, and it is also the unit of failure. An engine that lifts one block out of your page will lift it without the sentence three paragraphs up that qualified it.
What to change on Monday
Three things, in order of what the evidence actually supports.
Write blocks that survive being lifted. One complete claim per block, 150 to 300 words, the range inside the one bundle that was tested under control, and not a measured optimum. No “as we saw above,” no “in this section.” The entity name inside the block, not only in the H1. The figure with its date and unit inside the block, not in a table three screens down. Your reader carries context in their head. The retriever does not.
Make headings state answers, not topics. “Schema markup did not increase AI citations in a 1,885-page test” is retrievable. “About schema” is not. This is the cheapest change on the list and it is the one most sites skip.
Stop measuring whether you were cited, and start measuring which block was cited. Every AI visibility tracker hands you a URL. The useful work starts after that: open the page, find the passage that got used, and work out why that one and not the paragraph beside it. That is the only feedback loop that tells you what to write next, and no tool on the market will do it for you yet.
What to stop funding: additional markup, keyword density, and domain authority treated as a direct lever on citations. None of the four observations above, and neither controlled study, gives you a reason to keep paying for them.
The bottom line
The retrieval analysis is worth your time and it is not a leak. It is vendor research, disclosed by the author, run on a prompt about the vendor’s own category, and that goes on the record next to every number taken from it. Take the mechanism, which four researchers now agree on, and take the 150 to 300 word range, which sits inside a tested bundle rather than being read off a stream. Leave the scores, because there are three of them, from one page, on one shopping query, and the two claims most often quoted about them are not in the source at all.
The general rule survives the specific week. A number about how an AI system behaves will be repeated for years before anyone opens the primary document.
Be the one who opens it.
Corrections
Re-audited against every primary source on 16 September 2026. Three things changed.
- 01The author’s affiliation was missing. Yeşilyurt is a GEO researcher at Peec AI, the first post discloses it, and the piece now says so before quoting him.
- 02The 170-word blocks, the one-to-three blocks per source, the
should_fetchfield and the 150,000-line dump were cited to the second post. They are in the first. Only thesnippet::partsandsnippet::scoresfields and the three scores are in the second. - 03The 150 to 300 word range was described as tested under control. GEO-SFE tested its framework as a whole and recommends the range as a design principle; the 31% figure is cited from Liu et al. 2023 and the 23% figure carries no citation. The range is now described as part of a tested bundle, not a measured optimum, and the claim that it was “the only range measured under controlled conditions” is withdrawn.
Sources
- 01Metehan Yeşilyurt (Peec AI), “ChatGPT retrieval system LEAK alert: I found something huge in ChatGPT’s server-sent events over the weekend”, StartupTalky community, September 2026. https://community.startuptalky.com/discussions/post/chatgpt-retrieval-system-leak-alert-i-found-something-huge-in-EjJ2p6bMvfO2KSa
- 02Metehan Yeşilyurt (Peec AI), “ChatGPT retrieval system LEAK alert part 2: how ChatGPT ranks and scores”, StartupTalky community, September 2026. https://community.startuptalky.com/discussions/post/chatgpt-retrieval-system-leak-alert-part-2-how-chatgpt-ranks-and-scores-nICs0ezAMTuIt7U
- 03Metehan Yeşilyurt, “How I Reverse-Engineered ChatGPT’s Ranking Algorithm (And What It Means for Your Content)”, 20 August 2025. https://metehanai.substack.com/p/how-i-reverse-engineered-chatgpts
- 04Olivier de Segonzac (RESONEO), “Inside ChatGPT’s retrieval stack: The index, cache, and pages it actually reads”, Search Engine Land, 17 August 2026. https://searchengineland.com/chatgpt-retrieval-stack-index-cache-pages-485036
- 05Filippo Danesi, “Pages Get Retrieved. Chunks Get Cited.”, SERPsecrets, 15 July 2026. https://www.serp-secrets.com/blog/pages-retrieved-chunks-cited
- 06Junwei Yu, Mufeng Yang, Yepeng Ding, Hiroyuki Sato, “Structural Feature Engineering for Generative Engine Optimization: How Content Structure Shapes Citation Behavior”, arXiv preprint 2603.29979, submitted 31 March 2026. https://arxiv.org/abs/2603.29979
- 07Louise Linehan, “We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved.”, Ahrefs, 11 May 2026. https://ahrefs.com/blog/schema-ai-citations/
- 08Barry Schwartz, “Google passage ranking now live in US English search results”, Search Engine Land, 11 February 2021. https://searchengineland.com/google-passage-ranking-now-live-in-us-english-search-results-346034
- 09Peec AI, “Peec AI, the AI visibility monitoring company, assembles a team to decode how ChatGPT and other LLMs recommend brands”, GlobeNewswire, 26 June 2026. https://www.globenewswire.com/news-release/2026/06/26/3318325/0/en/Peec-AI-the-AI-visibility-monitoring-company-assembles-a-team-to-decode-how-ChatGPT-and-other-LLMs-recommend-brands.html
- 10Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang, “Lost in the Middle: How Language Models Use Long Contexts”, arXiv preprint 2307.03172, 2023. https://arxiv.org/abs/2307.03172
Questions
What is AI search optimization?
AI search optimization is the practice of making individual passages of a page retrievable and quotable by an answer engine, rather than making a whole page rank. The distinction is load-bearing: on the observations published so far, ChatGPT fetches a page, splits it, scores the parts separately, and puts one to three of those parts in front of the model. The rest of the page is never read. The unit of competition is the passage.
Does ChatGPT rank pages or passages?
Passages, on every published observation to date. Metehan Yeşilyurt, a GEO researcher at the AI visibility tracker Peec AI, read ChatGPT's server-sent events stream and found the page render cut into blocks of roughly 170 words, each carrying its own relevance score in fields named snippet::parts and snippet::scores, with the model shown one to three blocks per source. Independent network-traffic work relayed by SERPsecrets found citations binding to a specific sentence rather than to the page. Google shipped the same idea for classic search in February 2021 under the name passage ranking.
Was there a ChatGPT retrieval leak?
No. The data came from the server-sent events stream that ChatGPT uses to push its answer to the browser, which is visible to anyone with developer tools open. That makes it client-side telemetry rather than a breach, and the distinction matters in your favour: an observation anyone can repeat is stronger evidence than a secret you have to take on trust.
Does schema markup help AI citations?
The one controlled test says no. Ahrefs tracked 1,885 URLs that added JSON-LD between August 2025 and March 2026 against roughly 4,000 matched control pages, measured 30 days either side. ChatGPT citations moved 2.2% and Google AI Mode 2.4%, both indistinguishable from zero, while AI Overviews fell 4.6% against a platform-wide decline that hit the control group too. Schema markup for AI search is a symptom of a well-run site, not the cause of its citations.
How long should a passage be for AI search optimization?
The best-supported range is 150 to 300 words, and it is less settled than it sounds. The GEO-SFE preprint recommends that range as a design principle and tested its structural framework as a whole, measuring a 17.3% citation lift. It did not test paragraph length on its own. The 31% attention-degradation figure it prints is cited from a 2023 long-context study, and its 23% figure for chunks under 150 words carries no citation. Do not treat the 170-word figure from the ChatGPT stream as a target: it is one observer's reading of one product in one mode, and a second researcher measured roughly 200-character snippets in a different mode of the same product.
Which numbers from the ChatGPT retrieval analysis should not be trusted?
Two that reached us in secondhand summaries are not in either of the two posts they are attributed to, and the error is the summarisers', not the author's. The first is a claim that the top-scoring block is normalised to exactly 1.0000, which would make cross-query score comparison meaningless. The second is a claim that the findings were verified across 76 pre-indexed pages. We searched the full text of both posts and the author's own longer write-up and found neither. The scores that are published are three values from one page, on one shopping query, which the author himself flags as changing with A/B test variants.
This is analysis, not a verified outcome. It carries no verification badge and never will. The proof lives in the case files, where every figure is checked against the public record and the method is printed on the page.