← All case files
pending deployment software · US · ops

GitHub Copilot in the field: three company RCTs, 4,867 developers, 26% more completed tasks

Pooling three randomized controlled trials at Microsoft, Accenture and a Fortune 100 manufacturer, economists found that giving 4,867 software developers access to GitHub Copilot raised completed tasks by 26.08% (SE 10.3%), with the largest gains among junior developers.

MetricBeforeAfter
Number of completed tasks, treatment vs control (pooled across three RCTs, 4,867 developers) control group baseline +26.08% (SE: 10.3%)
Pull requests / commits / builds, Accenture experiment control group baseline pull requests +21.34% (SE 9.92%); commits +15.39% (SE 9.69%); builds +37.03% (SE 12.22%)
Output change by developer tenure (per MIT Sloan's writeup) senior developers +8% to +13% junior / recent-hire developers +27% to +39%

Verification status: CHECKING — handed to the checker, not yet graduated. NOT verified, NOT green. This page reports a peer-reviewed research finding; no green badge is claimed and none is implied. The headline number is precisely stated and independently published, but a vendor tie sits on it, and the page is built to make that impossible to miss.

The problem

The most-quoted number for AI coding productivity comes from a lab: a 2022 controlled experiment in which developers finished a single toy task 55.8% faster with GitHub Copilot (source). That figure is precise but it measures one standardized task run by the tool’s own maker, so it says little about whether the tool moves real output inside a working organization. To test that, economists turned to field data: “randomized controlled trials at Microsoft, Accenture, and an anonymous Fortune 100 company”, each “run by the companies as part of their ordinary course of business” (source).

What was built

Each experiment “provided a random subset of developers with access to an AI-based coding assistant suggesting intelligent code completions”, the assistant being “GitHub Copilot” (source). The trials were staggered across the three sites. The Microsoft experiment ran from “the week of September 2022” and involved “a sample size of 1,746 developers”; the Accenture experiment “started in the last week of July 2023 and included a number of Accenture offices located in Southeast Asia”, with “61.3% of the 320 developers assigned to the treatment group”; and the Fortune 100 experiment “started in October 2023” and “involved 3,054 developers” (source). Rather than a stopwatch on a toy task, the researchers measured real workflow output, taking “pull requests, which can be thought of as a unit of work for software developers”, along with commits and builds (source).

The outcome

The headline is a pooled estimate. “When data is combined across three experiments and 4,867 developers, our analysis reveals a 26.08% increase (SE: 10.3%) in completed tasks among developers using the AI tool” (source). MIT Sloan’s independent writeup states the same result in round terms, that “access to Copilot increased output — the number of completed weekly tasks — by 26%” (source). In the Accenture experiment, where the individual outcomes are reported, “the number of pull requests made by developers increases by 21.34% (SE: 9.92%), the number of commits increases by 15.39% (SE: 9.69%), and the number of builds increases by 37.03% (SE: 12.22%)” (source).

The gains were not evenly spread. MIT Sloan reports the effect concentrated in “recent hires and developers in more junior positions, who increased their output by 27% to 39%”, while “more senior developers saw productivity gains of 8% to 13%” (source). The paper’s own framing agrees: “less experienced developers had higher adoption rates and greater productivity gains” (source).

What the finding does and does not show

The result is a genuine field measurement, pre-registered as “AEARCTR-0014530” and peer-reviewed in Management Science, which is a far stronger design than a single lab task (source). The load-bearing limit is independence: the strongest source is the working paper itself, and while the lead economists are at Princeton, MIT and Wharton, two of the six authors are Microsoft employees and Copilot is a Microsoft and GitHub product, so this is not a fully arm’s-length evaluation of the vendor’s own tool (source). Two further limits are stated in the paper. The pooled estimate is noisy: the authors write that “though each experiment is noisy”, the signal appears only “when data is combined” (source). And a separate randomized trial by METR found the opposite sign for a different population, slowing experienced open-source developers by 19%, a reminder that the effect depends heavily on who is using the tool (source).

How this was verified

Retrieved live on 2026-09-04. Every figure is quoted verbatim from the MIT-hosted working-paper PDF (independent academic authors) and cross-checked against MIT Sloan’s independent writeup; each source was archived to the Wayback Machine and a copy of the paper saved in sources/. The paper is peer-reviewed in Management Science (INFORMS DOI 10.1287/mnsc.2025.00535) and pre-registered as AEARCTR-0014530. No figure was rounded or inferred; where the paper and the MIT Sloan writeup differ in precision (26.08% vs 26%), both are shown.

  • The same “juniors gain most” split appears in the BCG jagged-frontier RCT, where 758 consultants finished 12.2% more tasks with GPT-4 but the weakest performers gained the most.
  • The MIT ChatGPT writing trial is the other large randomized measure of a generative-AI productivity gain, 40% faster writing, and it too found the effect largest for the initially weaker workers.
  • The Bank of Korea’s 2026 issue note is the counterpoint: AI cut work time 3.8% yet the measured productivity gain was near zero, a caution against reading a task-count increase as an output increase.

Sources

  1. Kevin Zheyuan Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng, Tobias Salz · “The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers” (working paper) · February 2025 · https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdfTier 1 (independent academic authors; peer-reviewed in Management Science, INFORMS DOI 10.1287/mnsc.2025.00535; pre-registered AEARCTR-0014530). This is the weakest load-bearing source in one respect only, noted in the prose: two of the six authors are Microsoft employees and Copilot is a Microsoft/GitHub product.
  2. Dylan Walsh · “How generative AI affects highly skilled workers” · MIT Sloan (Ideas Made to Matter) · November 4, 2024 · https://mitsloan.mit.edu/ideas-made-to-matter/how-generative-ai-affects-highly-skilled-workersTier 2 (independent institutional writeup of the study).
  3. Sida Peng, Eirini Kalliamvakou, Peter Cihon, Mert Demirer · “The Impact of AI on Developer Productivity: Evidence from GitHub Copilot” · arXiv 2302.06590 · February 2023 · https://arxiv.org/abs/2302.06590Tier 3 (the 55.8% lab figure; authored by GitHub/Microsoft staff on their own product, cited here only as the contrast the field study set out to test). Checked live 2026-09-07.
  4. METR · “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity” · July 10, 2025 · https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/Tier 1 (independent nonprofit research organisation reporting its own randomized trial; the 19% slowdown). Checked live 2026-09-07.

GitHub Copilot (IDE code-completion assistant); developer telemetry (pull requests, commits, builds) analysed via a weighted instrumental-variables design

Verification record
Status
pending
Method
Retrieved live on 2026-09-04 from the MIT-hosted working-paper PDF (independent academic authors) and MIT Sloan's independent writeup; every figure quoted verbatim and archived to the Wayback Machine with a local copy in sources/. The paper is peer-reviewed in Management Science and pre-registered (AEARCTR-0014530). The page notes the one vendor tie: two of six authors are Microsoft employees and Copilot is a Microsoft/GitHub product.
Provider
GitHub Copilot (AI code-completion assistant)
Client
Software developers at Microsoft, Accenture, and an anonymous Fortune 100 electronics manufacturer (three RCT cohorts, 4,867 developers) · software
Disclosure
named
Questions this file answers
How much did GitHub Copilot raise developer output in the field experiments?

Pooling three randomized controlled trials at Microsoft, Accenture and a Fortune 100 manufacturer covering 4,867 developers, the researchers found a 26.08% increase (standard error 10.3%) in the number of completed tasks among developers given Copilot access.

Who ran the study and is it independent of GitHub?

The trials were run by the companies in their ordinary course of business and analysed by economists at Princeton, MIT, Wharton and Microsoft; it is peer-reviewed in Management Science and pre-registered as AEARCTR-0014530. Two of the six authors are Microsoft employees and Copilot is a Microsoft/GitHub product, so it is not fully arm's-length from the vendor, though the lead economists are independent.