# GitHub Copilot in the field: three company RCTs, 4,867 developers, 26% more completed tasks

> Pooling three randomized controlled trials at Microsoft, Accenture and a Fortune 100 manufacturer, economists found that giving 4,867 software developers access to GitHub Copilot raised completed tasks by 26.08% (SE 10.3%), with the largest gains among junior developers.

- Verification status: pending
- Case type: deployment
- Provider: GitHub Copilot (AI code-completion assistant) (https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf)
- Client: Software developers at Microsoft, Accenture, and an anonymous Fortune 100 electronics manufacturer (three RCT cohorts, 4,867 developers), software (named)
- Sector: software / US / ops
- Canonical URL: https://theinternetninja.com/stories/generative-ai-copilot-three-field-experiments-software-developers-26pct-tasks-2024/
- Source: The Internet Ninja (theinternetninja.com), independent verified-proof platform

## Outcomes

| Metric | Before | After |
| --- | --- | --- |
| Number of completed tasks, treatment vs control (pooled across three RCTs, 4,867 developers) | control group baseline | +26.08% (SE: 10.3%) |
| Pull requests / commits / builds, Accenture experiment | control group baseline | pull requests +21.34% (SE 9.92%); commits +15.39% (SE 9.69%); builds +37.03% (SE 12.22%) |
| Output change by developer tenure (per MIT Sloan's writeup) | senior developers +8% to +13% | junior / recent-hire developers +27% to +39% |

## Verification method

Retrieved live on 2026-09-04 from the MIT-hosted working-paper PDF (independent academic authors) and MIT Sloan's independent writeup; every figure quoted verbatim and archived to the Wayback Machine with a local copy in sources/. The paper is peer-reviewed in Management Science and pre-registered (AEARCTR-0014530). The page notes the one vendor tie: two of six authors are Microsoft employees and Copilot is a Microsoft/GitHub product.

## FAQ

**How much did GitHub Copilot raise developer output in the field experiments?**

Pooling three randomized controlled trials at Microsoft, Accenture and a Fortune 100 manufacturer covering 4,867 developers, the researchers found a 26.08% increase (standard error 10.3%) in the number of completed tasks among developers given Copilot access.

**Who ran the study and is it independent of GitHub?**

The trials were run by the companies in their ordinary course of business and analysed by economists at Princeton, MIT, Wharton and Microsoft; it is peer-reviewed in Management Science and pre-registered as AEARCTR-0014530. Two of the six authors are Microsoft employees and Copilot is a Microsoft/GitHub product, so it is not fully arm's-length from the vendor, though the lead economists are independent.

## Full case file

**Verification status: CHECKING — handed to the checker, not yet graduated. NOT verified,
NOT green.** This page reports a peer-reviewed research finding; no green badge is claimed and
none is implied. The headline number is precisely stated and independently published, but a
vendor tie sits on it, and the page is built to make that impossible to miss.

## The problem
The most-quoted number for AI coding productivity comes from a lab: a 2022 controlled
experiment in which developers finished a single toy task 55.8% faster with GitHub Copilot
([source](https://arxiv.org/abs/2302.06590)).
That figure is precise but it measures one standardized task run by the tool's own maker, so
it says little about whether the tool moves real output inside a working organization. To test
that, economists turned to field data: "randomized controlled trials at Microsoft, Accenture,
and an anonymous Fortune 100 company", each "run by the companies as part of their ordinary
course of business"
([source](https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf)).

## What was built
Each experiment "provided a random subset of developers with access to an AI-based coding
assistant suggesting intelligent code completions", the assistant being "GitHub Copilot"
([source](https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf)).
The trials were staggered across the three sites. The Microsoft experiment ran from "the week
of September 2022" and involved "a sample size of 1,746 developers"; the Accenture experiment
"started in the last week of July 2023 and included a number of Accenture offices located in
Southeast Asia", with "61.3% of the 320 developers assigned to the treatment group"; and the
Fortune 100 experiment "started in October 2023" and "involved 3,054 developers"
([source](https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf)).
Rather than a stopwatch on a toy task, the researchers measured real workflow output, taking
"pull requests, which can be thought of as a unit of work for software developers", along with
commits and builds
([source](https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf)).

## The outcome
The headline is a pooled estimate. "When data is combined across three experiments and 4,867
developers, our analysis reveals a <span class="kpi">26.08%</span> increase (SE: 10.3%) in
completed tasks among developers using the AI tool"
([source](https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf)).
MIT Sloan's independent writeup states the same result in round terms, that "access to Copilot
increased output — the number of completed weekly tasks — by <span class="kpi">26%</span>"
([source](https://mitsloan.mit.edu/ideas-made-to-matter/how-generative-ai-affects-highly-skilled-workers)).
In the Accenture experiment, where the individual outcomes are reported, "the number of pull
requests made by developers increases by <span class="kpi">21.34%</span> (SE: 9.92%), the
number of commits increases by <span class="kpi">15.39%</span> (SE: 9.69%), and the number of
builds increases by <span class="kpi">37.03%</span> (SE: 12.22%)"
([source](https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf)).

The gains were not evenly spread. MIT Sloan reports the effect concentrated in "recent hires
and developers in more junior positions, who increased their output by <span class="kpi">27%
to 39%</span>", while "more senior developers saw productivity gains of <span class="kpi">8%
to 13%</span>"
([source](https://mitsloan.mit.edu/ideas-made-to-matter/how-generative-ai-affects-highly-skilled-workers)).
The paper's own framing agrees: "less experienced developers had higher adoption rates and
greater productivity gains"
([source](https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf)).

## What the finding does and does not show
The result is a genuine field measurement, pre-registered as "AEARCTR-0014530" and peer-reviewed
in Management Science, which is a far stronger design than a single lab task
([source](https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf)).
The load-bearing limit is independence: the strongest source is the working paper itself, and
while the lead economists are at Princeton, MIT and Wharton, two of the six authors are
Microsoft employees and Copilot is a Microsoft and GitHub product, so this is not a fully
arm's-length evaluation of the vendor's own tool
([source](https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf)).
Two further limits are stated in the paper. The pooled estimate is noisy: the authors write
that "though each experiment is noisy", the signal appears only "when data is combined"
([source](https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf)).
And a separate randomized trial by METR found the opposite sign for a different population,
slowing experienced open-source developers by 19%, a reminder that the effect depends heavily on
who is using the tool
([source](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/)).

## How this was verified
Retrieved live on 2026-09-04. Every figure is quoted verbatim from the MIT-hosted working-paper
PDF (independent academic authors) and cross-checked against MIT Sloan's independent writeup;
each source was archived to the Wayback Machine and a copy of the paper saved in `sources/`. The
paper is peer-reviewed in Management Science (INFORMS DOI 10.1287/mnsc.2025.00535) and
pre-registered as AEARCTR-0014530. No figure was rounded or inferred; where the paper and the
MIT Sloan writeup differ in precision (26.08% vs 26%), both are shown.

## Related case files
- The same "juniors gain most" split appears in [the BCG jagged-frontier RCT](/stories/bcg-harvard-jagged-frontier-gpt4-consultants-12-2pct-more-tasks-2023/), where 758 consultants finished 12.2% more tasks with GPT-4 but the weakest performers gained the most.
- [The MIT ChatGPT writing trial](/stories/noy-zhang-mit-chatgpt-writing-rct-40pct-faster-18pct-quality-science-2023/) is the other large randomized measure of a generative-AI productivity gain, 40% faster writing, and it too found the effect largest for the initially weaker workers.
- [The Bank of Korea's 2026 issue note](/stories/bank-of-korea-2026-ai-adoption-cuts-work-time-3-8pct-productivity-gain-near-zero/) is the counterpoint: AI cut work time 3.8% yet the measured productivity gain was near zero, a caution against reading a task-count increase as an output increase.

## Sources
1. Kevin Zheyuan Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng, Tobias Salz · "The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers" (working paper) · February 2025 · https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf — **Tier 1** (independent academic authors; peer-reviewed in Management Science, INFORMS DOI 10.1287/mnsc.2025.00535; pre-registered AEARCTR-0014530). This is the weakest load-bearing source in one respect only, noted in the prose: two of the six authors are Microsoft employees and Copilot is a Microsoft/GitHub product.
2. Dylan Walsh · "How generative AI affects highly skilled workers" · MIT Sloan (Ideas Made to Matter) · November 4, 2024 · https://mitsloan.mit.edu/ideas-made-to-matter/how-generative-ai-affects-highly-skilled-workers — **Tier 2** (independent institutional writeup of the study).
3. Sida Peng, Eirini Kalliamvakou, Peter Cihon, Mert Demirer · "The Impact of AI on Developer Productivity: Evidence from GitHub Copilot" · arXiv 2302.06590 · February 2023 · https://arxiv.org/abs/2302.06590 — **Tier 3** (the 55.8% lab figure; authored by GitHub/Microsoft staff on their own product, cited here only as the contrast the field study set out to test). Checked live 2026-09-07.
4. METR · "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" · July 10, 2025 · https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ — **Tier 1** (independent nonprofit research organisation reporting its own randomized trial; the 19% slowdown). Checked live 2026-09-07.