# AI fracture flags raised residents' detection sensitivity from 84.7% to 91.3% at Charite Berlin, without losing specificity

> In a prospective study at Charite Universitaetsmedizin Berlin, showing residents the Gleamer BoneView AI fracture flags raised their fracture-detection sensitivity from 84.74% (human alone) to 91.28% (AI-assisted report) across 1,163 radiographs, with specificity holding at 97.11% then 97.36%. An independent meta-analysis of 26 BoneView studies (9,218 radiographs) reports the same direction: pooled clinician sensitivity 77% to 87%, specificity 88% to 92%.

- Verification status: pending
- Case type: deployment
- Provider: Gleamer BoneView, a CE-marked / FDA-cleared deep-learning fracture-detection assistant that flags and localises fractures on radiographs to support the reading clinician
- Client: Klinik fuer Radiologie, Charite Universitaetsmedizin Berlin (prospective single-center study, 1,163 exams in 735 patients), Healthcare - hospital radiology / emergency imaging (named)
- Sector: healthcare / DE / ops
- Canonical URL: https://theinternetninja.com/stories/gleamer-boneview-ai-fracture-detection-charite-berlin-84-7-to-91-3pct-sensitivity-2023/
- Source: The Internet Ninja (theinternetninja.com), independent verified-proof platform

## Outcomes

| Metric | Before | After |
| --- | --- | --- |
| Resident fracture-detection sensitivity, human-alone vs AI-assisted report |  |  |
| Pooled clinician sensitivity with vs without AI (independent meta-analysis, 26 studies) |  |  |

## Verification method

Two independent peer-reviewed sources carry the story, neither with a Gleamer author. The primary is the prospective Charite study (Life (Basel) 2023, 13(1):223, PMC9864518): the figures (84.74% / 86.92% / 91.28% sensitivity; 97.11% / 84.67% / 97.36% specificity; 1,163 exams / 735 patients / 367 fractures; 35 changes / 25 additional fractures) were quote-matched verbatim against the full text saved to sources/pmc-PMC9864518-charite.html and archived to web.archive.org (capture 20260909022256); it declares no external funding and names Gleamer only in the acknowledgments for technical support. The corroboration is an independent systematic review and meta-analysis (Annals of Medicine 2025, PMC12795274) of 26 studies / 9,218 radiographs: pooled clinician sensitivity 77%->87%, specificity 88%->92%, average paired gain +9.5%, quote-matched verbatim against sources/pmc-PMC12795274-meta.html (archived web.archive.org capture 20260909022506). The effect direction is further corroborated by a second independent reader study (residents' sensitivity 58%->77% with AI, PMC10969303) and BoneView's aggregate accuracy by a second independent diagnostic-test-accuracy meta-analysis (pooled sensitivity 0.90, PMC11452402). No confirmation was sought from Gleamer or Charite.

## FAQ

**How much did BoneView improve fracture detection at Charite Berlin?**

In a prospective study of 1,163 radiographs, residents' fracture-detection sensitivity rose from 84.74% reading alone to 91.28% when they were shown the BoneView AI flags, while specificity held at 97.11% then 97.36%. The AI alone read at 86.92% sensitivity.

**Is the BoneView fracture-detection gain independent of the vendor?**

The Charite study reports no external funding and names Gleamer only in its acknowledgments for technical support, and an independent meta-analysis of 26 BoneView studies (9,218 radiographs) found the same direction: pooled clinician sensitivity rose from 77% to 87% and specificity from 88% to 92% with AI assistance.

## Full case file

**Verification status: capped at corroborated. Not verified.**
Two independent peer-reviewed sources carry this file and neither is authored by the AI
vendor, which is unusually clean. What caps it below green is that the headline
84.74% to 91.28% figure is a single single-center study published in a lower-tier journal;
read the weak-source line below.

## The problem
Fracture reads on plain radiographs are done under time pressure, often by residents on
night shifts, and missed fractures are one of the more common diagnostic errors in
emergency imaging. The question the Charite study set out to answer was whether integrating
an AI fracture-detection assistant into the routine reading workflow would raise how many
fractures residents catch without making them flag healthy bones as broken. The residents'
reads were checked against a final diagnosis "made by a board-certified radiologist with
over eight years of experience, or if available, cross-sectional imaging"
([source](https://pmc.ncbi.nlm.nih.gov/articles/PMC9864518/)).

## What was built
The study integrated Gleamer BoneView, a deep-learning tool that flags and localises
fractures on radiographs, into the fracture-reading workflow at the radiology department of
Charite Universitaetsmedizin Berlin. Radiographs first read by two residents were then
re-read with the AI's results shown, and the change in the residents' performance was
measured. In total "<span class="kpi">1163</span> exams in
<span class="kpi">735</span> patients were included, with a total of
<span class="kpi">367</span> fractures (31.56%)"
([source](https://pmc.ncbi.nlm.nih.gov/articles/PMC9864518/)). The AI is an assist, not a
replacement: the flags are shown to the reading clinician, who makes the final call.

## The outcome
Showing the residents the AI flags raised how many fractures they caught. Per the paper,
"Pure human sensitivity was <span class="kpi">84.74%</span>, and AI sensitivity was
<span class="kpi">86.92%</span>. Thirty-five changes were made after showing AI results, 33
of which resulted in the correct diagnosis, resulting in 25 additionally found fractures.
This resulted in a sensitivity of <span class="kpi">91.28%</span> for the assisted report.
Specificity was <span class="kpi">97.11</span>, 84.67, and
<span class="kpi">97.36</span>%, respectively"
([source](https://pmc.ncbi.nlm.nih.gov/articles/PMC9864518/)). In plain terms, the
assisted read caught more fractures than either the residents or the AI alone, and it did so
"without a loss of specificity"
([source](https://pmc.ncbi.nlm.nih.gov/articles/PMC9864518/)).

That direction is not unique to one hospital. An independent systematic review and
meta-analysis pooled 26 BoneView studies covering 9,218 radiographs and found that "the
pooled sensitivity of clinicians increased from <span class="kpi">77%</span> (95% CI: 72-81)
to <span class="kpi">87%</span> (95% CI: 83-90) with AI assistance, while the pooled
specificity improved from <span class="kpi">88%</span> (95% CI: 85-90) to
<span class="kpi">92%</span> (95% CI: 89-94)"
([source](https://pmc.ncbi.nlm.nih.gov/articles/PMC12795274/)). Across those studies
"sensitivity improved by an average of <span class="kpi">9.5%</span> (95% CI: 6.8-12.1) with
AI assistance (p < 0.001)"
([source](https://pmc.ncbi.nlm.nih.gov/articles/PMC12795274/)). A separate independent
reader study at University Medicine Essen (no Gleamer author, funded by the German Research
Foundation) found the same direction in residents specifically: "Radiology residents'
sensitivity for fracture detection improved significantly with AI support
(<span class="kpi">58%</span> without AI vs. <span class="kpi">77%</span> with AI, p < 0.001)"
([source](https://pmc.ncbi.nlm.nih.gov/articles/PMC10969303/)). And a second independent
meta-analysis of commercial fracture-detection products reports that "the BoneView tool was
the most frequently studied and showed a pooled sensitivity of <span class="kpi">0.90</span>
(95% CI 0.85, 0.94) and specificity of <span class="kpi">0.89</span> (95% CI 0.87, 0.92)",
its authors declaring no competing interests
([source](https://pmc.ncbi.nlm.nih.gov/articles/PMC11452402/)).

**Weakest load-bearing source, named.** The specific 84.74% to 91.28% figure rests on a
single single-center study published in *Life (Basel)*, an MDPI journal whose peer review is
less selective than a top radiology title, so on its own that number is one hospital's
result in a lower-tier venue rather than a definitive effect size
([source](https://pmc.ncbi.nlm.nih.gov/articles/PMC9864518/)). What keeps the file honest is
that the meta-analysis is a separate, independent aggregate reaching the same conclusion, so
the direction and rough magnitude do not stand or fall with the Charite paper alone
([source](https://pmc.ncbi.nlm.nih.gov/articles/PMC12795274/)).

> **How this was verified.** Method: the Charite figures (84.74% / 86.92% / 91.28%
> sensitivity; 97.11 / 84.67 / 97.36% specificity; 1,163 exams / 735 patients / 367
> fractures; 35 changes / 25 additional fractures) were quote-matched verbatim against the
> peer-reviewed primary (Life (Basel) 2023, 13(1):223, PMC9864518), full text saved to
> sources/pmc-PMC9864518-charite.html and archived to web.archive.org
> (capture 20260909022256) - Tier 1, independent (no external funding; Gleamer named only in
> acknowledgments). The pooled figures (77%->87% sensitivity, 88%->92% specificity, +9.5%
> average gain, 26 studies / 9,218 radiographs) were quote-matched verbatim against the
> independent systematic review and meta-analysis (Annals of Medicine 2025, PMC12795274),
> saved to sources/pmc-PMC12795274-meta.html - Tier 1, independent aggregate, archived to
> web.archive.org (capture 20260909022506, content-verified). A round-2 exhaustion search
> added two further independent corroborations of the effect (residency reader study
> PMC10969303, 58%->77%; and a second diagnostic-test-accuracy meta-analysis PMC11452402,
> pooled BoneView sensitivity 0.90), both archived and saved to sources/. Date verified:
> 2026-09-09. No confirmation was sought from Gleamer or Charite:
> asking a vendor to confirm its own numbers is a testimonial, not an audit.

## Sources
1. Life (Basel), MDPI · *A Prospective Approach to Integration of AI Fracture Detection Software in Radiographs into Clinical Workflow* · Oppenheimer J, Lueken S, Hamm B, Niehues SM (Klinik fuer Radiologie, Charite Universitaetsmedizin Berlin) · 2023-01-13, 13(1):223 (PMC9864518) · https://pmc.ncbi.nlm.nih.gov/articles/PMC9864518/ — **Tier 1** (independent, peer-reviewed prospective single-center study; source of the 84.74%/86.92%/91.28% figures; declares no external funding, Gleamer named only in acknowledgments; note MDPI's *Life* is a lower-tier journal; full text saved to sources/pmc-PMC9864518-charite.html, archived web.archive.org/web/20260909022256).
2. Annals of Medicine · *Enhanced fracture detection on radiographs with AI assistance for clinicians: a systematic review and meta-analysis* · 2025 (PMC12795274) · https://pmc.ncbi.nlm.nih.gov/articles/PMC12795274/ — **Tier 1** (independent systematic review and meta-analysis of 26 BoneView studies / 9,218 radiographs; source of the pooled 77%->87% and 88%->92% figures and the +9.5% average gain; saved to sources/pmc-PMC12795274-meta.html; archived web.archive.org/web/20260909022506).
3. Diagnostics (Basel) · *AI-Assisted X-ray Fracture Detection in Residency Training: Evaluation in Pediatric and Adult Trauma Patients* · 2024 (PMC10969303) · https://pmc.ncbi.nlm.nih.gov/articles/PMC10969303/ — **Tier 2** (independent reader study, University Medicine Essen; DFG/UMEA funded, no Gleamer author, "the other authors declare no potential conflicts of interest"; corroborates the effect direction in residents, 58%->77% with AI; saved to sources/pmc-PMC10969303-residency-reader.html, archived web.archive.org/web/20260909025306).
4. Scientific Reports · *Artificial intelligence in commercial fracture detection products: a systematic review and meta-analysis of diagnostic test accuracy* · 2024 (PMC11452402) · https://pmc.ncbi.nlm.nih.gov/articles/PMC11452402/ — **Tier 1** (independent diagnostic-test-accuracy meta-analysis; "the authors declare no competing interests"; corroborates BoneView aggregate accuracy, pooled sensitivity 0.90 / specificity 0.89, though on AI-alone accuracy rather than the clinician-assisted before/after; saved to sources/pmc-PMC11452402-dta-meta.html, archived web.archive.org/web/20250215133539).

## Related case files
- [AI-ECG alerts in a Taiwan randomized trial](/stories/ai-ecg-alert-rct-taiwan-cuts-90-day-mortality-3-6-vs-4-3pct-nature-medicine-2024/) — the closest sibling in shape: an AI that flags risk to the treating clinician rather than replacing them, measured on what the human does better with the flag, here on ECGs instead of radiographs.
- [SAVE-O2, autonomous AI oxygen titration](/stories/save-o2-ai-autonomous-oxygen-titration-jama-85pct-vs-63pct-normoxemia-2026/) — the contrast on the autonomy axis: a clinical AI that acts in the loop on its own, where BoneView only ever assists the reader.
- [Boston Children's rare-disease reanalysis with a deep-research model](/stories/boston-childrens-manton-openai-o3-deep-research-rare-disease-reanalysis-18-diagnoses-376-cases-nejm-ai-2026/) — another diagnostic-support result from a named hospital in the peer-reviewed record, where the AI surfaces candidates and clinicians make the call.

## Path to green
The direction and rough magnitude of the effect are corroborated by four independent
peer-reviewed sources with no Gleamer author: the prospective Charite study, a 26-study
meta-analysis of clinician-assisted reads, an independent residency reader study (58% to 77%
with AI), and a second diagnostic-test-accuracy meta-analysis (pooled BoneView sensitivity
0.90). A round-2 record-exhaustion search found no independent source that restates the
specific 84.74% to 91.28% Charite figure, which is a single-center number and irreproducible
by nature, and no second meta-analysis of the exact clinician-assisted 77% to 87% pooled
figures. The meta-analysis archive gap has been closed with a working, content-verified
snapshot. The public record is therefore exhausted for lifting the exact headline figures to
two independent sources; what the search did do is strengthen the corroborated band, since
the effect direction and BoneView's aggregate accuracy each now rest on more than one
independent study. No confirmation should be sought from Gleamer or Charite. On this record the
file is closed at the corroborated cap.