A World Bank RCT in Nigeria: a six-week GPT-4 tutoring program lifted learning by 0.31 standard deviations
In a randomized controlled trial in nine public schools in Benin City, Nigeria, first-year senior secondary students who used Microsoft Copilot (powered by GPT-4) as an after-school English tutor for six weeks scored 0.31 standard deviations higher on a composite assessment and 0.23 standard deviations higher on English than a control group. The World Bank team that ran the evaluation put the gains at 1.5 to 2 years of 'business-as-usual' schooling and among the most cost-effective interventions it has studied. Every figure is quoted verbatim from the World Bank's own working paper and blog and corroborated by independent reporting.
| Metric | Before | After |
|---|---|---|
| A six-week RCT deploying Microsoft Copilot (GPT-4) as a virtual tutor produced a 0.31 standard-deviation gain on a composite assessment and 0.23 standard deviations on English, the main outcome (World Bank Working Paper 11125, May 2025; Tier 1 primary; corroborated by independent press) | ||
| The World Bank evaluators put the gains at 1.5 to 2 years of 'business-as-usual' schooling and said the program outperformed roughly 80% of rigorously evaluated interventions in the developing world (World Bank, primary) | ||
| The trial ran in nine public schools in Benin City, Edo State, with about 800 first-year senior secondary students in after-school sessions between June and July 2024 (independent press + World Bank blog) | ||
| The largest effects were for female students and for students with higher initial academic performance, with benefit across the baseline-ability distribution (World Bank Working Paper 11125) | ||
Verification status: PENDING — the headline figures (a 0.31 standard-deviation composite gain and a 0.238 standard-deviation English gain from a six-week randomized controlled trial) are quoted verbatim from the World Bank’s own working paper, the independent evaluator of the program, whose full text and results tables were retrieved for this pass, and corroborated by the World Bank’s education blog and two independent outlets. Limits are flagged below: this is a single World Bank working paper not yet peer-reviewed or independently replicated, press and the paper report the English effect slightly differently (0.24 vs 0.238 standard deviations), and the “two years of schooling” figure is a cost-effectiveness equivalence, not a grade-level claim.
The problem
Nigeria has one of the world’s largest populations of children who are in school but not learning to the level their grade expects, and one-to-one tutoring — the intervention with the strongest evidence base — is expensive to deliver at scale. The World Bank team framed the question directly: whether a large language model could act as a virtual tutor in a low-resource classroom. The study “evaluates the impact of a program leveraging large language models for virtual tutoring in secondary education in Nigeria” (source). It was written up as World Bank Policy Research Working Paper 11125, “From Chalkboards to Chatbots,” in May 2025.
What was built
The program put a general-purpose chatbot in front of students as a supervised tutor rather than a homework shortcut. “Using a randomized controlled trial, the program deployed Microsoft Copilot (powered by GPT-4) to support first-year senior secondary students in English language learning over six weeks” (source). The rollout was small and concrete: it was “implemented over six weeks in nine public schools in Benin City” (source), where “in mid-2024, 800 first-year senior secondary students attended after-school English classes” (source). The World Bank’s own account places the intervention “between June and July 2024” (source), with students working in teacher-supervised computer-lab sessions rather than unsupervised on their phones.
The outcome
The measured effect was large for an education intervention. The paper reports that “the intervention demonstrated a significant improvement of 0.31 standard deviation on an assessment that included English topics aligned with the Nigerian curriculum, knowledge of artificial intelligence and digital skills. The effect on English, the main outcome of interest, was of 0.23 standard deviations” (source). The results tables carry the precision behind those headline numbers: “the treatment effect on the total score (weighted) is 0.31 standard deviation (SE = 0.068),” with the English effect at “0.238 σ, SE = 0.068” and both English and AI-skills effects “statistically significant at the 1% level,” across models where “the number of observations ranges between 636 and 654” (source). The gain also carried into an exam the program had not been tailored to: treatment “had a positive and significant impact on the third term exam score, with an effect size of 0.206 standard deviation (SE = 0.067),” an exam whose “content evaluated … was broader than the one covered during the six weeks” (source). The World Bank’s education team described the same result in rounder terms: “the learning improvements were striking—about 0.3 standard deviations” (source), and independent reporting recorded that “students in the treatment group scored 0.31 standard deviations higher on the final standardized assessment than their control-group peers” (source). The evaluators translated the gain into schooling: “cost-effectiveness analysis revealed substantial learning gains, equating to 1.5 to 2 years of ‘business-as-usual’ schooling, situating the intervention among some of the most cost-effective programs to improve learning outcomes” (source). On the World Bank’s own blog the team added the benchmark: “this is equivalent to nearly two years of typical learning in just six weeks. When we compared these results to a database of education interventions studied through randomized controlled trials in the developing world, our program outperformed 80% of them, including some of the most cost-effective strategies” (source). That ranking is worked through in the paper itself against a published review that “found a median effect of 0.10 standard deviations in overall test scores and 0.14 in reading (Evans and Yuan, 2022),” so that “the results found in this study are situated at least at the 80th percentile of all RCTs,” and “even when considering only RCTs that had between 500 and 1000 participants, the results are still higher than 80% of the other studies,” while “considering only the effects on language outcomes, the results are near the 70th percentile” (source). The gains were not uniform: the paper notes “while the program benefits students across the baseline ability distribution, the largest effects are for female students, and those with higher initial academic performance” (source).
The evidence behind the claim
The load-bearing source here is the evaluator, not the vendor. The number is stated by an independent World Bank research team that designed and ran the randomized controlled trial, not by Microsoft or OpenAI, and it is corroborated by two outlets with no stake in the tool: ICTworks, which recorded that “the magnitude of observed learning gains was remarkable, with effect sizes approximating 0.3 standard deviations” (source), and Devdiscourse, which reported “the intervention generated learning equivalent to two years of traditional schooling in Nigeria” (source). What is settled is what the trial measured over six weeks against a control group; what it does not settle is durability, scale beyond nine schools, or whether the effect holds outside a supervised computer-lab setting.
The weakest link in this record is no longer document access — the full working paper, including its results tables, was retrieved for this pass from the World Bank’s open-knowledge repository, so the effect sizes, standard errors and sample counts above are quoted from the tables rather than the abstract. The honest limit now is that this is a single World Bank working paper: rigorously executed, but not yet peer-reviewed in a journal and not yet independently replicated by a second team. The “outperformed 80% of interventions” ranking is the paper’s own benchmarking against one published review (Evans and Yuan, 2022), not an external audit of the comparison. Two smaller conflicts are shown rather than merged: press coverage reports the English effect as 0.24 standard deviations while the paper’s tables state 0.238, and the “two years of schooling” figure is the evaluators’ cost-effectiveness equivalence (a learning-adjusted-years-of-schooling translation), not a claim that a student advanced two grade levels in six weeks.
How this was verified
Method: every load-bearing figure was quoted verbatim from the World Bank’s Policy Research Working Paper 11125 (“From Chalkboards to Chatbots,” May 2025). The paper’s abstract was read via the RePEc/IDEAS mirror (sources/repec-wp11125.html), and its full text, including the results tables, was retrieved from the World Bank open-knowledge repository (bitstream cd9cca5b-2ffd-4be1-8313-0f4d2ed6ef4f, canonical landing at documents.worldbank.org) and saved to sources/wp11125-fulltext.txt, then grep-verified for each quoted figure (Tier 1, the independent evaluator). The effect sizes and schooling equivalence were corroborated against the World Bank’s education blog (January 9, 2025), ICTworks (January 23, 2025), and Devdiscourse (May 25, 2025), each saved to sources/ and grep-verified. Wayback archiving was unreachable from the research box (HTTP 000 / 523 / timeout), so the snapshots are the local sources/ copies and the archive gap is recorded for the checker. No confirmation was sought from the World Bank, Microsoft, or the schools; only the public record was used. Date of verification: September 4, 2026.
Related case files
- A US AI-tutor RCT where the measured gains matched a non-AI comparison — the counterweight to this result
- An earlier education RCT of an AI chatbot with a small but real measured effect on enrollment
- The other side of generative AI in education: an ed-tech incumbent whose model was disrupted by ChatGPT
Sources
- World Bank (De Simone, Tiberti, Barron Rodriguez, Manolio, Mosuro, Dikoru) · “From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria,” Policy Research Working Paper 11125 (full text incl. results tables) · May 2025 · https://documents.worldbank.org/en/publication/documents-reports/documentdetail/099548105192529324 — Tier 1 (primary; the independent evaluator’s own paper)
- World Bank · “From Chalkboards to Chatbots,” Policy Research Working Paper 11125 (abstract, via RePEc/IDEAS mirror) · May 2025 · https://ideas.repec.org/p/wbk/wbrwps/11125.html — Tier 1 (primary; abstract of the same paper)
- World Bank education blog · “From chalkboards to chatbots: Transforming learning in Nigeria, one prompt at a time” · January 9, 2025 · https://blogs.worldbank.org/en/education/From-chalkboards-to-chatbots-Transforming-learning-in-Nigeria — Tier 2 (first-party institutional, same evaluator)
- ICTworks (Wayan Vota) · “Generative AI Now Proven to Advance Learning Outcomes in Nigeria” · January 23, 2025 · https://www.ictworks.org/genai-advance-learning-outcomes/ — Tier 2 (independent press)
- Devdiscourse · “GPT-4 Tutoring in Nigeria Boosts English Scores, Offers Scalable, Cost-Effective Model” · May 25, 2025 · https://www.devdiscourse.com/article/education/3419251-gpt-4-tutoring-in-nigeria-boosts-english-scores-offers-scalable-cost-effective-model — Tier 2 (independent press)
Microsoft Copilot (powered by GPT-4) used as a supervised after-school virtual tutor, with teachers present and prompts designed to promote reasoning rather than shortcuts; delivered in school computer labs over twelve sessions
- Status
- pending
- Method
- Every load-bearing figure is quoted verbatim from the World Bank's own Policy Research Working Paper 11125, 'From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria' (May 2025), whose abstract was fetched this session via the RePEc/IDEAS mirror (ideas.repec.org/p/wbk/wbrwps/11125.html) and saved to sources/repec-wp11125.html (Tier 1, the independent evaluator; the paper's official home at openknowledge.worldbank.org served a JavaScript app / HTTP 403 to fetchers). The effect sizes and the schooling-equivalence framing are corroborated by the World Bank's own education blog, 'From chalkboards to chatbots: Transforming learning in Nigeria, one prompt at a time' (January 9, 2025, sources/wb-blog-transforming.html), and by two independent outlets: ICTworks (Wayan Vota, January 23, 2025, sources/ictworks.html) and Devdiscourse (May 25, 2025, sources/devdiscourse.html). Press reports the English effect as 0.24 SD while the paper's own abstract states 0.23 SD; both are quoted and the conflict is shown, not merged. Wayback archiving was unreachable from the research box this session (HTTP 523 / timeout), so snapshots are the local sources/ copies and the archive gap is flagged for the checker. No confirmation was sought from the World Bank, Microsoft, or the schools. Date of verification: September 4, 2026.
- Provider
- Microsoft Copilot (powered by GPT-4), used as an after-school virtual tutor; the program was designed and independently evaluated by a World Bank research team (Policy Research Working Paper 11125)
- Client
- Nine public senior secondary schools in Benin City, Edo State, Nigeria; evaluated by the World Bank · education
- Disclosure
- named
What did the World Bank study in Nigeria find?
A randomized controlled trial of a six-week after-school program that used Microsoft Copilot (powered by GPT-4) as a virtual tutor found a 0.31 standard-deviation improvement on a composite assessment and a 0.23 standard-deviation effect on English, the main outcome of interest, for first-year senior secondary students in Benin City, Nigeria.
How large is a 0.31 standard-deviation gain?
The World Bank evaluators put the measured learning gains at 1.5 to 2 years of 'business-as-usual' schooling and said the program outperformed roughly 80 percent of rigorously evaluated education interventions studied through randomized controlled trials in the developing world. The equivalence is the evaluators' cost-effectiveness framing, not a claim that students literally advanced two grade levels.
How was the trial run?
The program ran in nine public schools in Benin City, Edo State, Nigeria, with about 800 first-year senior secondary students, in after-school computer-lab sessions between June and July 2024. Students were randomly assigned to a treatment or control group, and the treatment group used Microsoft Copilot under teacher supervision with prompts designed to promote reasoning.