AI fracture flags raised residents' detection sensitivity from 84.7% to 91.3% at Charite Berlin, without losing specificity
In a prospective study at Charite Universitaetsmedizin Berlin, showing residents the Gleamer BoneView AI fracture flags raised their fracture-detection sensitivity from 84.74% (human alone) to 91.28% (AI-assisted report) across 1,163 radiographs, with specificity holding at 97.11% then 97.36%. An independent meta-analysis of 26 BoneView studies (9,218 radiographs) reports the same direction: pooled clinician sensitivity 77% to 87%, specificity 88% to 92%.
| Metric | Before | After |
|---|---|---|
| Resident fracture-detection sensitivity, human-alone vs AI-assisted report | ||
| Pooled clinician sensitivity with vs without AI (independent meta-analysis, 26 studies) | ||
Verification status: capped at corroborated. Not verified. Two independent peer-reviewed sources carry this file and neither is authored by the AI vendor, which is unusually clean. What caps it below green is that the headline 84.74% to 91.28% figure is a single single-center study published in a lower-tier journal; read the weak-source line below.
The problem
Fracture reads on plain radiographs are done under time pressure, often by residents on night shifts, and missed fractures are one of the more common diagnostic errors in emergency imaging. The question the Charite study set out to answer was whether integrating an AI fracture-detection assistant into the routine reading workflow would raise how many fractures residents catch without making them flag healthy bones as broken. The residents’ reads were checked against a final diagnosis “made by a board-certified radiologist with over eight years of experience, or if available, cross-sectional imaging” [source].
What was built
The study integrated Gleamer BoneView, a deep-learning tool that flags and localises fractures on radiographs, into the fracture-reading workflow at the radiology department of Charite Universitaetsmedizin Berlin. Radiographs first read by two residents were then re-read with the AI’s results shown, and the change in the residents’ performance was measured. In total “1163 exams in 735 patients were included, with a total of 367 fractures (31.56%)” [source]. The AI is an assist, not a replacement: the flags are shown to the reading clinician, who makes the final call.
The outcome
Showing the residents the AI flags raised how many fractures they caught. Per the paper, “Pure human sensitivity was 84.74%, and AI sensitivity was 86.92%. Thirty-five changes were made after showing AI results, 33 of which resulted in the correct diagnosis, resulting in 25 additionally found fractures. This resulted in a sensitivity of 91.28% for the assisted report. Specificity was 97.11, 84.67, and 97.36%, respectively” [source]. In plain terms, the assisted read caught more fractures than either the residents or the AI alone, and it did so “without a loss of specificity” ([source]).
That direction is not unique to one hospital. An independent systematic review and meta-analysis pooled 26 BoneView studies covering 9,218 radiographs and found that “the pooled sensitivity of clinicians increased from 77% (95% CI: 72-81) to 87% (95% CI: 83-90) with AI assistance, while the pooled specificity improved from 88% (95% CI: 85-90) to 92% (95% CI: 89-94)” [source]. Across those studies “sensitivity improved by an average of 9.5% (95% CI: 6.8-12.1) with AI assistance (p < 0.001)” [source]. A separate independent reader study at University Medicine Essen (no Gleamer author, funded by the German Research Foundation) found the same direction in residents specifically: “Radiology residents’ sensitivity for fracture detection improved significantly with AI support (58% without AI vs. 77% with AI, p < 0.001)” [source]. And a second independent meta-analysis of commercial fracture-detection products reports that “the BoneView tool was the most frequently studied and showed a pooled sensitivity of 0.90 (95% CI 0.85, 0.94) and specificity of 0.89 (95% CI 0.87, 0.92)”, its authors declaring no competing interests [source].
Weakest load-bearing source, named. The specific 84.74% to 91.28% figure rests on a single single-center study published in Life (Basel), an MDPI journal whose peer review is less selective than a top radiology title, so on its own that number is one hospital’s result in a lower-tier venue rather than a definitive effect size [source]. What keeps the file honest is that the meta-analysis is a separate, independent aggregate reaching the same conclusion, so the direction and rough magnitude do not stand or fall with the Charite paper alone ([source]).
How this was verified. Method: the Charite figures (84.74% / 86.92% / 91.28% sensitivity; 97.11 / 84.67 / 97.36% specificity; 1,163 exams / 735 patients / 367 fractures; 35 changes / 25 additional fractures) were quote-matched verbatim against the peer-reviewed primary (Life (Basel) 2023, 13(1):223, PMC9864518), full text saved to sources/pmc-PMC9864518-charite.html and archived to web.archive.org (capture 20260909022256) - Tier 1, independent (no external funding; Gleamer named only in acknowledgments). The pooled figures (77%->87% sensitivity, 88%->92% specificity, +9.5% average gain, 26 studies / 9,218 radiographs) were quote-matched verbatim against the independent systematic review and meta-analysis (Annals of Medicine 2025, PMC12795274), saved to sources/pmc-PMC12795274-meta.html - Tier 1, independent aggregate, archived to web.archive.org (capture 20260909022506, content-verified). A round-2 exhaustion search added two further independent corroborations of the effect (residency reader study PMC10969303, 58%->77%; and a second diagnostic-test-accuracy meta-analysis PMC11452402, pooled BoneView sensitivity 0.90), both archived and saved to sources/. Date verified: 2026-09-09. No confirmation was sought from Gleamer or Charite: asking a vendor to confirm its own numbers is a testimonial, not an audit.
Sources
- 01Life (Basel), MDPI · A Prospective Approach to Integration of AI Fracture Detection Software in Radiographs into Clinical Workflow · Oppenheimer J, Lueken S, Hamm B, Niehues SM (Klinik fuer Radiologie, Charite Universitaetsmedizin Berlin) · 2023-01-13, 13(1):223 (PMC9864518) · https://pmc.ncbi.nlm.nih.gov/articles/PMC9864518/ — Tier 1 (independent, peer-reviewed prospective single-center study; source of the 84.74%/86.92%/91.28% figures; declares no external funding, Gleamer named only in acknowledgments; note MDPI’s Life is a lower-tier journal; full text saved to sources/pmc-PMC9864518-charite.html, archived web.archive.org/web/20260909022256).
- 02Annals of Medicine · Enhanced fracture detection on radiographs with AI assistance for clinicians: a systematic review and meta-analysis · 2025 (PMC12795274) · https://pmc.ncbi.nlm.nih.gov/articles/PMC12795274/ — Tier 1 (independent systematic review and meta-analysis of 26 BoneView studies / 9,218 radiographs; source of the pooled 77%->87% and 88%->92% figures and the +9.5% average gain; saved to sources/pmc-PMC12795274-meta.html; archived web.archive.org/web/20260909022506).
- 03Diagnostics (Basel) · AI-Assisted X-ray Fracture Detection in Residency Training: Evaluation in Pediatric and Adult Trauma Patients · 2024 (PMC10969303) · https://pmc.ncbi.nlm.nih.gov/articles/PMC10969303/ — Tier 2 (independent reader study, University Medicine Essen; DFG/UMEA funded, no Gleamer author, “the other authors declare no potential conflicts of interest”; corroborates the effect direction in residents, 58%->77% with AI; saved to sources/pmc-PMC10969303-residency-reader.html, archived web.archive.org/web/20260909025306).
- 04Scientific Reports · Artificial intelligence in commercial fracture detection products: a systematic review and meta-analysis of diagnostic test accuracy · 2024 (PMC11452402) · https://pmc.ncbi.nlm.nih.gov/articles/PMC11452402/ — Tier 1 (independent diagnostic-test-accuracy meta-analysis; “the authors declare no competing interests”; corroborates BoneView aggregate accuracy, pooled sensitivity 0.90 / specificity 0.89, though on AI-alone accuracy rather than the clinician-assisted before/after; saved to sources/pmc-PMC11452402-dta-meta.html, archived web.archive.org/web/20250215133539).
Related case files
- AI-ECG alerts in a Taiwan randomized trial — the closest sibling in shape: an AI that flags risk to the treating clinician rather than replacing them, measured on what the human does better with the flag, here on ECGs instead of radiographs.
- SAVE-O2, autonomous AI oxygen titration — the contrast on the autonomy axis: a clinical AI that acts in the loop on its own, where BoneView only ever assists the reader.
- Boston Children’s rare-disease reanalysis with a deep-research model — another diagnostic-support result from a named hospital in the peer-reviewed record, where the AI surfaces candidates and clinicians make the call.
Path to green
The direction and rough magnitude of the effect are corroborated by four independent peer-reviewed sources with no Gleamer author: the prospective Charite study, a 26-study meta-analysis of clinician-assisted reads, an independent residency reader study (58% to 77% with AI), and a second diagnostic-test-accuracy meta-analysis (pooled BoneView sensitivity 0.90). A round-2 record-exhaustion search found no independent source that restates the specific 84.74% to 91.28% Charite figure, which is a single-center number and irreproducible by nature, and no second meta-analysis of the exact clinician-assisted 77% to 87% pooled figures. The meta-analysis archive gap has been closed with a working, content-verified snapshot. The public record is therefore exhausted for lifting the exact headline figures to two independent sources; what the search did do is strengthen the corroborated band, since the effect direction and BoneView’s aggregate accuracy each now rest on more than one independent study. No confirmation should be sought from Gleamer or Charite. On this record the file is closed at the corroborated cap.
Gleamer BoneView - deep-learning fracture detection and localisation on radiographsHuman-in-the-loop: the AI flags are shown to the reading clinician, who makes the final callProspective integration into the routine fracture-reading workflow of a university hospital
Verification record
- Status
- corroborated
- Method
- Two independent peer-reviewed sources carry the story, neither with a Gleamer author. The primary is the prospective Charite study (Life (Basel) 2023, 13(1):223, PMC9864518): the figures (84.74% / 86.92% / 91.28% sensitivity; 97.11% / 84.67% / 97.36% specificity; 1,163 exams / 735 patients / 367 fractures; 35 changes / 25 additional fractures) were quote-matched verbatim against the full text saved to sources/pmc-PMC9864518-charite.html and archived to web.archive.org (capture 20260909022256); it declares no external funding and names Gleamer only in the acknowledgments for technical support. The corroboration is an independent systematic review and meta-analysis (Annals of Medicine 2025, PMC12795274) of 26 studies / 9,218 radiographs: pooled clinician sensitivity 77%->87%, specificity 88%->92%, average paired gain +9.5%, quote-matched verbatim against sources/pmc-PMC12795274-meta.html (archived web.archive.org capture 20260909022506). The effect direction is further corroborated by a second independent reader study (residents' sensitivity 58%->77% with AI, PMC10969303) and BoneView's aggregate accuracy by a second independent diagnostic-test-accuracy meta-analysis (pooled sensitivity 0.90, PMC11452402). No confirmation was sought from Gleamer or Charite.
- Confidence
- Corroborated (0.70 to 0.95): independently sourced, below the verification bar
- Ceiling
- The 84.7% to 91.3% headline rests on a single-center prospective study in Life (Basel), and a record-exhaustion search found no independent source restating the exact clinician-assisted figures, so the ceiling sits at corroborated.
- Provider
- Gleamer BoneView, a CE-marked / FDA-cleared deep-learning fracture-detection assistant that flags and localises fractures on radiographs to support the reading clinician
- Client
- Klinik fuer Radiologie, Charite Universitaetsmedizin Berlin (prospective single-center study, 1,163 exams in 735 patients) · Healthcare - hospital radiology / emergency imaging
- Disclosure
- named
Questions this file answers
How much did BoneView improve fracture detection at Charite Berlin?
In a prospective study of 1,163 radiographs, residents' fracture-detection sensitivity rose from 84.74% reading alone to 91.28% when they were shown the BoneView AI flags, while specificity held at 97.11% then 97.36%. The AI alone read at 86.92% sensitivity.
Is the BoneView fracture-detection gain independent of the vendor?
The Charite study reports no external funding and names Gleamer only in its acknowledgments for technical support, and an independent meta-analysis of 26 BoneView studies (9,218 radiographs) found the same direction: pooled clinician sensitivity rose from 77% to 87% and specificity from 88% to 92% with AI assistance.