ai coding agent: what the measured record shows about developer productivity (2026)
2026-09-11
The benchmarks rank the model. The field trials measured real developers, and the range runs from 26% more work to 19% slower. Here is what the evidence supports.
Built on verified case files. The argument below leans on evidence The Internet Ninja validated against the public record and published in full, method included.
- GitHub Copilot in the field: three company RCTs, 4,867 developers, 26% more completed tasks
- The 'jagged frontier' experiment: GPT-4 let 758 BCG consultants finish 12.2% more tasks 25.1% faster, yet made them 19% less likely to be right on a task outside AI's reach
- ChatGPT cut Stack Overflow's question volume by about 25% in six months: the peer-reviewed measure of an AI-disrupted knowledge platform, and the 28% layoff that followed
Search “ai coding agent” and the results are rankings, benchmark charts and lists of the free tools. Every one answers the same question: which agent scores highest. Almost none answers the question a team lead actually has, which is whether a real team ships more work once the tool is installed.
Those are different questions, and the gap between them is where money gets wasted. A model that tops a coding benchmark can still leave your developers no faster, or slower. The only way to know is a controlled trial on working developers, and a handful now exist.
The largest found a real gain. Pooling three randomized controlled trials at Microsoft, Accenture and a Fortune 100 manufacturer, economists measured 4,867 developers and found GitHub Copilot raised completed tasks by 26.08%, with a standard error of 10.3% (source). That is the number a buyer’s guide should lead with, and rarely does.
What an ai coding agent is
An ai coding agent is a software tool built on a large language model that reads, writes and edits code inside a developer’s workflow, from inline completion up to multi-step tasks it plans and executes with limited supervision. The measured evidence to date is heaviest on the assistant end of that range, where the tool suggests and completes code a developer accepts or rejects.
That distinction matters for reading any benchmark. Most of the controlled evidence measures code-completion assistants like Copilot, not the newer autonomous agents that open pull requests on their own. Treat a result on one as a result on the other and you are guessing.
Do ai coding agent benchmarks predict real productivity?
Not reliably. A benchmark scores a model on a fixed task set under controlled conditions. It tells you how the model ranks, not how your developers will fare on your codebase, and the field trials show those two can diverge sharply.
The clearest demonstration is the “jagged frontier” experiment. In a pre-registered study published in Organization Science, 758 Boston Consulting Group consultants using GPT-4 completed 12.2% more tasks 25.1% faster on 18 tasks inside the model’s reach (source). On one task deliberately chosen to sit just outside that reach, the consultants using AI were 19% less likely to be correct.
Same tool, same people, opposite result, decided entirely by which side of the frontier the task fell on. A benchmark cannot tell you where your work sits relative to that edge, which is the one thing you need to know.
How much faster do ai coding agents make developers?
The honest answer is a range, not a number, and the range crosses zero. The Copilot field trials found a 26.08% rise in completed tasks, concentrated among junior developers, whose output rose 27% to 39% against 8% to 13% for seniors (source).
Then read the counterweight. METR ran a randomized controlled trial with 16 experienced open-source developers working on their own mature repositories. AI tools made them 19% slower, even though the same developers believed the tools had sped them up (source).
Both results are real and neither cancels the other. The gain shows up for developers early on a task or a codebase they do not yet know well. The slowdown shows up for experts on code they know cold, where reviewing a suggestion costs more than writing the line. The tool did not change; the frontier did.
The proof: what TIN verified
TIN’s advantage on this beat is that it does not run on impressions. Each of the studies above is a verified case file, every figure re-checked against the primary source and archived.
- The Copilot field experiments. Three company RCTs, 4,867 developers, 26.08% more completed tasks, with the Accenture cohort also showing pull requests up 21.34% and builds up 37.03%. Read the verified case file. The one vendor tie is noted on the page: two of the six authors are Microsoft employees.
- The jagged frontier. GPT-4 lifted in-frontier work and degraded out-of-frontier work in the same experiment. Read the verified case file.
- The second-order effect. After ChatGPT, activity on Stack Overflow fell about 25% within six months, largest for the most widely used programming languages, with no significant change in post quality (source). The public knowledge these agents were trained on is the same knowledge developers stopped contributing. Read the verified case file.
ai coding agents comparison: what the field trials actually compared
A fair comparison names its axis. These trials did not rank tools against each other; they compared developers with a tool against developers without one. That is the comparison a benchmark leaderboard skips, and it is the one that decides whether a licence pays for itself.
| Study | Population | Design | Measured effect |
|---|---|---|---|
| Copilot field experiments | 4,867 developers | 3 RCTs, pooled | +26.08% completed tasks |
| Jagged frontier | 758 consultants | Pre-registered field experiment | +12.2% in-frontier, 19% less likely correct out-of-frontier |
| METR | 16 experienced OSS developers | RCT on own repositories | 19% slower |
Three controlled studies, three different populations, and effects from plus 26% to minus 19%. No leaderboard would have predicted that spread, because the spread is not about the model. It is about who is using it and on what.
The bottom line
Stop asking which ai coding agent ranks highest and start asking where your work sits relative to the tool’s frontier. The measured record says the gain is real for developers new to a task and can invert to a loss for experts on code they know cold. A ranking cannot locate that line for you, and a vendor will not.
The transferable principle outlasts any one tool: when a technology helps unevenly, the average benchmark score is the least useful number about it. What you need is the shape of the curve, and only a trial on people like yours draws that.
Sources
- Cui, Demirer, Jaffe, Musolff, Peng, Salz, “The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers”, working paper, February 2025 (peer-reviewed in Management Science, DOI 10.1287/mnsc.2025.00535). https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf
- METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”, 10 July 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- Dell’Acqua et al., “Navigating the Jagged Technological Frontier”, Organization Science (INFORMS), Vol. 37 No. 2, published online 11 March 2026. https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838
- del Rio-Chanona, Laurentsyeva, Wachs, “Large language models reduce public knowledge sharing on online Q&A platforms”, PNAS Nexus 3(9):pgae400, 11 September 2024. https://academic.oup.com/pnasnexus/article/3/9/pgae400/7754871
Questions
Do ai coding agents make developers faster?
Sometimes, and the range is wide. Three company field trials of 4,867 developers found GitHub Copilot raised completed tasks 26.08%, but a separate randomized trial of experienced open-source developers found AI made them 19% slower on their own repositories.
Do ai coding agent benchmarks predict real productivity?
Not reliably. A benchmark scores the model on fixed tasks, but the field experiments show the effect depends on where your work sits relative to the tool's capability, which a leaderboard does not measure.
What is the strongest evidence that an ai coding agent helps?
The strongest independent evidence is a pooled analysis of three randomized controlled trials covering 4,867 software developers, which found a 26.08% rise in completed tasks, with junior developers gaining most.
Sources
- Cui, Demirer, Jaffe, Musolff, Peng, Salz (working paper, peer-reviewed in Management Science), The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers , 2025-02-01
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity , 2025-07-10
- Organization Science (INFORMS), Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality , 2026-03-11
- PNAS Nexus, Large language models reduce public knowledge sharing on online Q&A platforms , 2024-09-11
This is analysis, not a verified outcome. It carries no verification badge and never will. The proof lives in the case files, where every figure is checked against the public record and the method is printed on the page.