← All case files
verified deployment healthcare · US · ops

An OpenAI reasoning model reanalyzed 376 unsolved rare-disease cases at Boston Children's. Experts confirmed 18 new diagnoses.

In an NEJM AI study, Boston Children's Manton Center, Harvard and OpenAI ran 376 previously unsolved pediatric rare-disease cases through the OpenAI o3 Deep Research model; after expert ACMG/AMP review and CLIA-certified lab confirmation, physicians established 18 new diagnoses, a 4.8% additional yield on cases specialists had already failed to solve.

MetricBeforeAfter
Additional diagnostic yield from AI-assisted reanalysis of already-unsolved cases
Neurodevelopmental cohort
Neuromuscular cohort
Sudden unexpected death in pediatrics cohort
Early psychosis cohort
Diagnoses that were 'rediscoveries' absent from the reviewed record

The problem

Even with modern genome sequencing, “roughly half” of people with rare diseases “remain undiagnosed after extensive testing and specialist review,” because the answer can require sifting “thousands to millions of possible genetic variants, fragmented clinical records, and rapidly changing scientific literature” (source). An inconclusive genetic test is not permanent: as new gene-disease links and variant reclassifications accumulate, an old unsolved case “can become newly interpretable,” so institutions inherit a growing backlog of genomes to keep in sync with a moving knowledge base (source). The Manton Center for Orphan Disease Research at Boston Children’s Hospital works with “over 3,500 individuals globally,” routinely re-screening their genomes, yet those screenings “often do not turn up any new answers” (source).

What was built

Researchers from the Manton Center, Harvard University and OpenAI built an expert-led reanalysis workflow around the OpenAI o3 Deep Research reasoning model, run last year when “it was then the most powerful system available” (source). For each case the team assembled a de-identified packet of standardized Human Phenotype Ontology terms, occasional clinician notes, and “a filtered variant table” capturing each variant’s rarity, predicted protein effect and ClinVar classification, then asked the model “to propose the most plausible molecular explanation and to show its work” (source). The model was an “explanation-first reasoning layer on top of existing genomic pipelines,” not a ranked-gene returner (source). Crucially, “the model did not diagnose any patient or make any clinical decision”; a finding counted as a diagnosis “only after qualified experts reviewed the evidence, the variant was classified as pathogenic or likely pathogenic, a CLIA-certified laboratory confirmed it, and the clinical team returned the result to the family” (source).

The outcome

The headline result: 18 of 376 previously unsolved cases newly diagnosed, an additional diagnostic yield of 4.8% (95% CI 2.9 to 7.5), split across four cohorts: 10 of 100 neurodevelopmental (10.0%), 4 of 61 neuromuscular (6.6%), 2 of 200 sudden unexpected death in pediatrics (1.0%) and 2 of 15 early psychosis (13.3%); 7 of 18 were rediscoveries.

The team ran “376 previously unsolved genetic cases” through the workflow and, after review and clinical confirmation, physicians “established diagnoses in 18 cases,” an “additional diagnostic yield of 4.8% after earlier analysis by specialists” (source). Lead researcher Catherine Brownstein framed why a small percentage matters here: “It got almost 5% new diagnoses, which doesn’t sound like a lot, but considering how many times these had already been analyzed, that’s a huge number, and each one means an answer for a family” (source). The 18 diagnoses split across four cohorts: neurodevelopmental “10 patients,” neuromuscular “four patients,” “two children who had died suddenly” and “two patients with early childhood psychosis illnesses” (source). The peer-reviewed NEJM AI abstract states the full result directly: “new local diagnoses were made in 10 of 100 rare disease neurodevelopmental cases (10.0%…), 4 of 61 neuromuscular cases (6.6%…), 2 of 200 cases of sudden unexpected death in pediatrics (1.0%…), and 2 of 15 early psychosis cases (13.3%…) for an overall diagnostic yield of 18 of 376 (4.8%, [CI, 2.9 to 7.5])” (source). OpenAI’s own results table gives the same denominators and rates (source), and independent trade coverage reports the same cohort yields of “10%,” “6.6%,” “13.3%” and “1%” (source).

What the number does not mean

Seven of the 18 diagnoses were “rediscoveries,” cases where “a treatment team in one location had identified a patient’s specific diagnosis but had not shared it with researchers around the world,” so the AI’s contribution there was surfacing an answer that already existed elsewhere, not creating new knowledge (source). The study was retrospective, “the cohorts were heterogeneous, and reviewers were not blinded to model confidence,” and the researchers “did not measure time saved, cost, clinician effort, false-positive workload, or changes in care” (source). Independent experts interviewed by NBC News called the result meaningful while cautioning that “LLM results still require rigorous human review” before any diagnosis is reported to a patient (source).

Weakest load-bearing source

The headline figures now rest on the peer-reviewed primary itself: the NEJM AI abstract (DOI 10.1056/aics2501343, PMID 42529462), read verbatim from the MEDLINE record mirrored by Europe PMC, states the 18-of-376 yield, the 4.8% overall rate with a 95% confidence interval, and the full four-cohort table with denominators (source). Two vendor-independent press reports corroborate the same numbers: NBC News confirms the 376 cases, the 18 diagnoses and the near-5% yield (source), and CLP Mag, a clinical-lab trade publication, independently prints the exact “4.8% diagnostic yield” and the four cohort rates (source). The caveat worth naming plainly is no longer a sourcing gap but a conflict of interest and a study design: “OpenAI helped financially support the project,” and OpenAI both supplied the model and co-authored the paper, so its own applied-AI page stays a Tier-3 interested source used only where the primary already carries the fact (source). The result is also retrospective, and the authors themselves call for prospective multi-center evaluation, which does not yet exist (source).

How this was verified Method: every figure was traced verbatim to the peer-reviewed primary plus independent secondary coverage, all fetched this session. The NEJM AI abstract (DOI 10.1056/aics2501343, PMID 42529462), read from the MEDLINE record mirrored by Europe PMC, is the Tier-1 origin of the 18-of-376 yield, the 4.8% rate with its confidence interval, the four-cohort table with denominators, the seven rediscoveries and the a-priori ACMG/AMP + CLIA + return-to-family diagnosis definition. NBC News (“AI helped diagnose 18 children whose rare diseases had stumped doctors,” 2026-06-18) and CLP Mag / Clinical Lab Products (2026-06-19) independently corroborate the same numbers; OpenAI’s own applied-AI write-up (2026-06-18, Tier 3) is used only where the primary already carries the fact. The NEJM AI full text at ai.nejm.org stayed behind a Cloudflare anti-bot wall, but the abstract surface carries every cited figure. NBC News, CLP Mag and the OpenAI page were archived on the Wayback Machine; the Europe PMC Tier-1 surface is a stable, re-fetchable public EBI/MEDLINE URL, captured verbatim in the case file’s source record, with Wayback re-archival flagged for the checker. Date: 2026-08-25.

Path to green

The critical figures are now anchored to the peer-reviewed NEJM AI primary (its MEDLINE-indexed abstract), which independently confirms 376 cases, 18 diagnoses, the 4.8% overall yield with its confidence interval and the four-cohort table with denominators, so the earlier Tier-1 gap on the numbers is closed. What remains is generalisability, not sourcing: the study is retrospective, and the record would need what the authors themselves call for, a prospective multi-center evaluation with predefined end points comparing LLM-assisted reanalysis with standard practice on diagnostic yield, time to a candidate, clinician effort, false-positive burden and cost. No such replication exists on the public record today, and the standing conflict of interest, that OpenAI both built the model and helped fund the study, would only be removed by a fully vendor-independent measurement. Those set the ceiling; the headline figures themselves are corroborated to the primary.

Sources

  1. NBC News · Jared Perlo, “AI helped diagnose 18 children whose rare diseases had stumped doctors” · 2026-06-18 · https://www.nbcnews.com/tech/innovation/ai-boston-childrens-hospital-diagnose-rare-diseases-kids-openai-rcna350387. Tier 2 (archived: https://web.archive.org/web/20260821190554/https://www.nbcnews.com/tech/innovation/ai-boston-childrens-hospital-diagnose-rare-diseases-kids-openai-rcna350387)
  2. CLP / Clinical Lab Products · “AI Model Improves Rare Disease Diagnosis Rates, NEJM AI Study Shows” · 2026-06-19 · https://clpmag.com/diagnostic-technologies/molecular-diagnostics/ai-reasoning-model-boosts-rare-disease-diagnosis-yield/. Tier 2 (archived: https://web.archive.org/web/20260825110248/https://clpmag.com/diagnostic-technologies/molecular-diagnostics/ai-reasoning-model-boosts-rare-disease-diagnosis-yield/)
  3. OpenAI · “Using AI to help physicians diagnose rare genetic diseases affecting children” · 2026-06-18 · https://openai.com/index/diagnose-rare-childhood-diseases/. Tier 3 (provider’s own; OpenAI helped fund the study) (archived: https://web.archive.org/web/20260803010802/https://openai.com/index/diagnose-rare-childhood-diseases/)
  4. Today’s Clinical Lab · Janette Wider, “AI Identifies New Rare Disease Diagnoses in Previously Unsolved Pediatric Cases” · 2026 · https://www.clinicallab.com/ai-identifies-new-rare-disease-diagnoses-in-previously-unsolved-pediatric-cases-28714. Tier 2 (secondary, itself sourced to NBC News; Wayback capture returned origin error 520, not archived)
  5. NEJM AI (primary, peer-reviewed), abstract via Europe PMC / MEDLINE · Jaech A, Cheatham M, et al., “LLM-Assisted Reanalysis of Unsolved Rare Disease Genomes Increases Diagnostic Yield,” NEJM AI 3(7), DOI 10.1056/aics2501343, PMID 42529462 · 2026-06-25 · https://europepmc.org/article/MED/42529462. Tier 1 (peer-reviewed abstract fetched verbatim this session; carries the 18/376, 4.8% CI 2.9 to 7.5 yield, the four-cohort table with denominators, the seven rediscoveries and the ACMG/AMP + CLIA diagnosis definition. Also at the EBI REST core record https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=DOI:10.1056/AIcs2501343&resultType=core&format=json . The NEJM AI full text at https://ai.nejm.org/doi/full/10.1056/AIcs2501343 stayed behind a Cloudflare anti-bot wall; Wayback capture of the Europe PMC surface did not complete this session and is flagged for re-archival.)

OpenAI o3 Deep Research reasoning modelHuman Phenotype Ontology (HPO)filtered variant table (rarity, protein effect, ClinVar)ACMG/AMP variant classificationCLIA-certified confirmatory testing

Verification record
Status
verified
Method
Every figure was traced verbatim to the peer-reviewed NEJM AI abstract (DOI 10.1056/aics2501343, PMID 42529462), fetched this session from the MEDLINE record mirrored by Europe PMC, which carries the 18/376 (4.8%, CI 2.9-7.5) yield, the four-cohort table with denominators, the seven rediscoveries and the ACMG/AMP + CLIA diagnosis definition. Two vendor-independent press reports corroborate the same numbers: NBC News (2026-06-18) and Clinical Lab Products / CLP Mag (2026-06-19). OpenAI's own Tier-3 write-up is used only where the primary already carries the fact. The NEJM AI full text stayed behind an anti-bot wall; the abstract surface supplies every cited figure. Press sources archived on the Wayback Machine; the Europe PMC Tier-1 surface is a stable re-fetchable EBI/MEDLINE URL with Wayback re-archival flagged.
Provider
OpenAI o3 Deep Research reasoning model
Client
Boston Children's Hospital, Manton Center for Orphan Disease Research (with Harvard University) · Pediatric rare-disease genomics / academic medicine
Disclosure
named
Questions this file answers
Did OpenAI's AI diagnose rare diseases at Boston Children's?

No. The o3 Deep Research model proposed the most plausible molecular explanation for expert review; a case counted as a diagnosis only after qualified experts classified the variant as pathogenic or likely pathogenic, a CLIA-certified lab confirmed it, and the clinical team returned the result to the family.

How many new diagnoses did the AI-assisted reanalysis find?

Across 376 previously unsolved cases, physicians established 18 new diagnoses, a 4.8% additional yield (95% CI 2.9 to 7.5). Seven of the 18 were rediscoveries of answers already established elsewhere but absent from the local research record.