# Tutor CoPilot: a Stanford randomized trial found AI real-time coaching made tutors more effective, with the biggest gains for the weakest tutors

> In the first randomized controlled trial of a human-AI system in live tutoring, Stanford researchers found that K-12 students whose tutors had Tutor CoPilot were 4 percentage points more likely to master topics, with students of lower-rated tutors gaining 9 percentage points, at an estimated cost of $20 per tutor per year.

- Verification status: pending
- Case type: deployment
- Provider: Tutor CoPilot (Stanford EduNLP Lab), deployed with the tutoring company FEV Tutor
- Client: FEV Tutor / a large U.S. school district in the South (K-12 tutoring), education (named)
- Sector: education / US / cross
- Canonical URL: https://theinternetninja.com/stories/tutor-copilot-stanford-fev-rct-4pp-mastery-9pp-lower-rated-tutors-2024/
- Source: The Internet Ninja (theinternetninja.com), independent verified-proof platform

## Outcomes

| Metric | Before | After |
| --- | --- | --- |
| Likelihood of mastering a topic in a tutoring session | Tutors without AI assistance | +4 percentage points (p<0.01) |
| Topic mastery for students of lower-rated tutors | Lower-rated tutors without AI assistance | +9 percentage points |
| Estimated cost of the tool | Human-only expert coaching (expensive to scale) | $20 per tutor per year |

## Verification method

Independent public-record verification. The effect sizes (4 p.p., 9 p.p.), the trial scale (900 tutors, 1,800 K-12 students) and the $20 per-tutor cost are quoted verbatim from the primary Stanford working paper (arXiv 2410.03017, Tier 1), retrieved live this session and archived; each figure is independently corroborated by education-trade press K-12 Dive (Tier 2), which also supplies the named deployment partner (FEV Tutor) and the district's region, and the two headline effects (4 p.p. and 9 p.p.) are corroborated by a second independent outlet, The 74 (Tier 2). No author or vendor was contacted; a research finding is verified against the primary artifact, not its authors.

## Full case file

## The problem
Training novice tutors with expert guidance raises the quality of tutoring, but that guidance is expensive and hard to scale, which "creates significant barriers to improving education quality at scale" and "disproportionately harms students from under-served communities, who stand to gain the most from high-quality education" ([source](https://arxiv.org/abs/2410.03017)). The open question was whether a language model could stand in for that expert coaching in a live session, and whether that actually raised what students learned rather than merely promising to ([source](https://arxiv.org/abs/2410.03017)).

## What was built
This is a research finding about a human-AI system, not a vendor deployment. Stanford researchers built "Tutor CoPilot, a novel Human-AI approach that leverages a model of expert thinking to provide expert-like guidance to tutors as they tutor" ([source](https://arxiv.org/abs/2410.03017)). The tool coaches the tutor in real time during a session; the tutor, not the AI, works with the student. To test it, the team ran "the first randomized controlled trial of a Human-AI system in live tutoring, involving 900 tutors and 1,800 K-12 students from historically under-served communities" ([source](https://arxiv.org/abs/2410.03017)). The independent trade outlet K-12 Dive reports the deployment partner and setting the paper anonymises: "Stanford partnered with tutoring company FEV Tutor to pilot the tool's implementation" and the "1,800 elementary and secondary school students" came "from a large school district in the South" ([source](https://www.k12dive.com/news/ai-tutor-effectiveness-stanford-university/728980/)).

## The outcome
Following a preregistered analysis plan, the trial found a modest but statistically significant gain: "students working with tutors that have access to Tutor CoPilot are 4 percentage points (p.p.) more likely to master topics (p<0.01)" ([source](https://arxiv.org/abs/2410.03017)). K-12 Dive reports the same figure independently, noting students whose tutors used the tool "were 4 percentage points more likely to progress through math tutoring session assessments successfully" ([source](https://www.k12dive.com/news/ai-tutor-effectiveness-stanford-university/728980/)). A second independent outlet, The 74, reports the same effect, that students "who worked with AI-assisted tutors were four percentage points more likely to master the topic after a given session than those in a control group whose tutors didn't work with AI" ([source](https://www.the74million.org/article/study-ai-assisted-tutoring-boosts-students-math-skills/)).

The effect was not evenly distributed, and this is the study's central finding: "students of lower-rated tutors experienced the greatest benefit, improving mastery by 9 p.p." ([source](https://arxiv.org/abs/2410.03017)). K-12 Dive corroborates it, reporting that students of lower-rated tutors "increased their math proficiency up to 9 percentage points on average" ([source](https://www.k12dive.com/news/ai-tutor-effectiveness-stanford-university/728980/)). The 74 independently reports the same larger effect, that students working with lower-rated tutors "saw their performance jump more than twice as much, by nine percentage points" ([source](https://www.the74million.org/article/study-ai-assisted-tutoring-boosts-students-math-skills/)). The tool was cheap relative to human coaching: the researchers "find that Tutor CoPilot costs only $20 per-tutor annually" ([source](https://arxiv.org/abs/2410.03017)), a figure K-12 Dive reports as "$20 per tutor annually, based on the tutors' usage during the study" ([source](https://www.k12dive.com/news/ai-tutor-effectiveness-stanford-university/728980/)).

The researchers also looked at how tutors changed their behaviour. Analysing more than half a million messages, they report: "We analyze 550,000+ messages using classifiers to identify pedagogical strategies, and find that tutors with access to Tutor CoPilot are more likely to use high-quality strategies to foster student understanding (e.g., asking guiding questions) and less likely to give away the answer to the student" ([source](https://arxiv.org/abs/2410.03017)). The same paper flags a limit the tool has not solved, reporting that tutors "flag issues in Tutor CoPilot, such as generating suggestions that are not grade-level appropriate" ([source](https://arxiv.org/abs/2410.03017)).

The weakest load-bearing point here is that these effects rest on a single working paper that had not, at the time of writing, completed peer review in a journal; the primary source is the authors' own arXiv preprint, and the two fully independent outlets fetched, K-12 Dive and The 74, each reproduce the same headline numbers rather than measuring anything separately ([source](https://www.the74million.org/article/study-ai-assisted-tutoring-boosts-students-math-skills/)). The named deployment partner and the district's region also come only from that press report, because the paper itself anonymises the site ([source](https://www.k12dive.com/news/ai-tutor-effectiveness-stanford-university/728980/)). This page therefore documents what a single, independently designed and preregistered RCT found, not a replicated or journal-published result.

## How this was verified
Method: on 1 September 2026 the primary Stanford working paper (arXiv 2410.03017, "Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise", v2 dated 26 January 2025) was retrieved live, and its abstract figures — the 4 percentage-point mastery effect, the 9 percentage-point effect for students of lower-rated tutors, the 900-tutor / 1,800-student scale, the $20 per-tutor cost and the 550,000-plus message analysis — were copied verbatim. Each figure was independently corroborated against K-12 Dive's report (published 7 October 2024), which also supplied the named partner (FEV Tutor) and the district's region, details the paper anonymises. The two headline effects (4 p.p. and 9 p.p.) were further corroborated against a second independent outlet, The 74 (published 7 October 2024). Both sources were archived to the Wayback Machine and saved locally to `sources/`. This is independent academic research evaluating a human-AI system, so it is verified against the primary paper itself; no author, institution or vendor was contacted, in line with TIN's independent-audit rule. Limit: this is one preregistered but unreplicated working paper, so the record supports what a single well-designed RCT found, not a replicated finding.

## Related case files
- [A World Bank RCT in Nigeria where six weeks of generative-AI tutoring raised learning by about 0.3 standard deviations, another independent, preregistered read on AI's measured effect on real students](/stories/world-bank-nigeria-genai-tutoring-rct-0-3sd-edo-state-2025/)
- [An NBER RCT of the Khanmigo AI tutor in Tennessee where math gains matched non-AI tutoring, a contrasting independent result on whether AI tutoring beats the status quo](/stories/nber-khanmigo-ai-tutor-tennessee-rct-math-gains-match-non-ai-2026/)
- [The BCG-Harvard "jagged frontier" experiment, where AI lifted consultants most for the lower performers, the same pattern of AI compressing the gap between weaker and stronger workers seen here](/stories/bcg-harvard-jagged-frontier-gpt4-consultants-12-2pct-more-tasks-2023/)
- [The Bank of Korea's finding that AI adoption cut work time but the measured productivity gain was near zero, an independent counterweight on the distance between AI's promise and its realised effect](/stories/bank-of-korea-2026-ai-adoption-cuts-work-time-3-8pct-productivity-gain-near-zero/)

## Sources
1. Rose E. Wang, Ana T. Ribeiro, Carly D. Robinson, Susanna Loeb, Dora Demszky (Stanford University) · "Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise" · arXiv 2410.03017 (v2, 2025-01-26; v1 2024-10-03) · https://arxiv.org/abs/2410.03017 — **Tier 1** (primary preregistered working paper; source of every figure; single study, not yet peer-reviewed in a journal)
2. K-12 Dive (Kara Arundel) · "How AI can improve tutor effectiveness" · 2024-10-07 · https://www.k12dive.com/news/ai-tutor-effectiveness-stanford-university/728980/ — **Tier 2** (independent education-trade press; corroborates the 4 p.p. and 9 p.p. effects, the 900/1,800 scale and the $20 cost, and names FEV Tutor and the district's region)
3. The 74 (Greg Toppo) · "Study: AI-Assisted Tutoring Boosts Students' Math Skills" · 2024-10-07 · https://www.the74million.org/article/study-ai-assisted-tutoring-boosts-students-math-skills/ — **Tier 2** (independent education news; second independent outlet corroborating the 4 p.p. mastery effect and the 9 p.p. effect for students of lower-rated tutors)