An MIT randomized trial in Science: ChatGPT cut professional writing time 40% and raised graded quality 18%
In a preregistered randomized controlled trial published in Science in July 2023, MIT researchers Shakked Noy and Whitney Zhang gave 453 college-educated professionals occupation-specific writing tasks and randomly gave half access to ChatGPT. Access cut the time to finish the tasks by 40 percent and raised output quality, as graded by independent evaluators, by 18 percent, while inequality between workers fell because weaker performers gained the most. Every figure is quoted verbatim from MIT's own release about the peer-reviewed paper and corroborated by the authors' pre-peer-review working paper.
| Metric | Before | After |
|---|---|---|
| A preregistered RCT of 453 professionals found ChatGPT cut task completion time by 40% and raised independently graded output quality by 18% (Science, July 2023; Tier 1 peer-reviewed; corroborated by the authors' working paper) | ||
| Performance inequality between workers decreased: lower-graded workers benefited most, so ChatGPT compressed the productivity distribution (Science / MIT working paper) | ||
| The tasks were two 20-to-30-minute occupation-specific writing assignments (cover letters, sensitive emails, short analytical pieces) the researchers called broadly representative of the professionals' real work (MIT) | ||
Verification status: CHECKING — the headline figures (a 40 percent reduction in task time and an 18 percent rise in independently graded quality across 453 professionals) are quoted verbatim from MIT’s own release about the peer-reviewed Science paper, and corroborated by the authors’ pre-peer-review working paper, which reports the same result in standard-deviation units. Limits are flagged below: the published article body itself bot-blocked fetchers this session, so the exact percentages are quoted from MIT’s institutional release rather than the Science text, and the trial measured short controlled writing tasks in early 2023, not sustained job performance.
The problem
By early 2023 the question about generative AI had shifted from whether it could write to whether it actually made real workers faster or better, and most of the loud claims came from vendors with a product to sell. Two MIT economics PhD students, Shakked Noy and Whitney Zhang, set out to answer it with an experiment instead of an anecdote. Their aim, as MIT put it, was “to study generative AI’s effect on worker productivity” directly, using a controlled trial rather than testimonials (source). The work was written up in the peer-reviewed journal Science, where “the study … appears … in open-access form” (source).
What was built
This was not a product deployment but a designed experiment. The researchers “assign occupation-specific, incentivized writing tasks to … college-educated professionals, and randomly expose half of them to ChatGPT” (source). In practice, “the researchers gave 453 college-educated marketers, grant writers, consultants, data analysts, human resource professionals, and managers two writing tasks specific to their occupation” (source). The assignments were short and realistic: “the 20- to 30-minute tasks included writing cover letters for grant applications, emails about organizational restructuring,” and similar pieces, and “the researchers say the tasks were broadly representative of assignments such professionals see in their real jobs” (source). Output quality was not self-reported: it was “measured by independent evaluators” (source).
The outcome
The effect was large and went in both directions that matter, speed and quality at once. According to MIT, access to “the assistive chatbot ChatGPT decreased the time it took workers to complete the tasks by 40 percent, and output quality, as measured by independent evaluators, rose by 18 percent” (source). The authors’ own working paper states the same finding in the units economists use for effect sizes: “our results show that ChatGPT substantially raises average productivity: time taken decreases by 0.8 SDs and output quality rises by 0.4 SDs” (source). The distributional result is the one that has drawn the most attention: “the data also showed that performance inequality between workers decreased, meaning workers who received a lower grade in the first task benefitted more from using ChatGPT for the second task” (source). The working paper puts the mechanism plainly: “inequality between workers decreases, as ChatGPT compresses the productivity distribution by benefiting low-ability workers more,” and it “mostly substitutes for worker effort rather than complementing worker skills, and restructures tasks towards idea-generation and editing and away from rough-drafting” (source).
The evidence behind the claim
The load-bearing source is independent published science, not a vendor. The trial was preregistered, the tasks were incentivized, and quality was graded by independent evaluators rather than by the participants or by OpenAI, and the result was published in Science after peer review (source). What is settled is what the experiment measured: on short, occupation-specific writing tasks in a controlled online setting in early 2023, ChatGPT made this sample of professionals substantially faster and their output somewhat better, and it helped the weakest performers most. What it does not settle is whether the same gains hold over sustained real-world work, across every occupation, or for later or different AI tools.
The weakest link in this record is source access, and it is worth naming precisely. The exact published percentages (40 percent faster, 18 percent higher quality, on 453 participants) are quoted here from MIT’s own institutional news release about the paper, because the peer-reviewed Science article page and the SSRN and Ovid mirrors returned bot-blocks (HTTP 403 and 402) to fetchers this session, so the article body itself could not be read directly. That release is MIT reporting a peer-reviewed result by its own researchers, not an independent outlet, so it is treated as strong but not primary. It is backed here by a source that was fetched directly: the authors’ own pre-peer-review working paper, which reports the identical result in standard-deviation units (0.8 SD faster, 0.4 SD higher quality) on a slightly smaller pre-revision sample of 444 participants (source). The move from 444 participants and standard-deviation units in the working paper to 453 and percentage terms in the published Science version is a revision between drafts, shown here rather than merged, not a conflict in the finding.
How this was verified
Method: the headline figures were quoted verbatim from MIT’s news release, “Study finds ChatGPT boosts worker productivity for some writing tasks” (MIT News, July 14, 2023), which was fetched this session, saved to sources/mit-news-chatgpt-productivity.html, and grep-verified for 453 college-educated, by 40 percent and quality rose by 18 percent (Tier 2, MIT reporting its own peer-reviewed study). The peer-reviewed article of record is Noy and Zhang, “Experimental evidence on the productivity effects of generative artificial intelligence,” Science 381:187-192 (July 2023, DOI 10.1126/science.adh2586); the article page and its SSRN and Ovid mirrors served bot-blocks (HTTP 403/402) to fetchers this session, so it is cited as the primary of record but its body was not read directly this pass. The authors’ pre-peer-review working paper (MIT Department of Economics, March 2, 2023) was fetched and saved to sources/noy-zhang-mit-econ-working-paper.pdf and independently corroborates the finding in standard-deviation units on 444 participants (Tier 1 primary artifact). Wayback archiving was unreachable from the research box (timeout), so the snapshots are the local sources/ copies and the archive gap is recorded for the checker. No confirmation was sought from OpenAI or the researchers; only the public record was used. Date of verification: September 5, 2026.
Related case files
- A sibling randomized trial where GPT-4 lifted consultants’ task completion but only inside the tool’s “jagged frontier”
- A field RCT of generative AI in customer support that found the same inequality-compression pattern — least-experienced agents gained most
- The counter-case: a controlled trial where AI coding tools made experienced open-source developers slower, not faster
Sources
- Shakked Noy and Whitney Zhang · “Experimental evidence on the productivity effects of generative artificial intelligence” · Science 381:187-192 · July 13, 2023 · https://www.science.org/doi/10.1126/science.adh2586 — Tier 1 (primary, peer-reviewed; article body bot-blocked to fetchers this session, cited as the publication of record)
- MIT News · “Study finds ChatGPT boosts worker productivity for some writing tasks” · July 14, 2023 · https://news.mit.edu/2023/study-finds-chatgpt-boosts-worker-productivity-writing-0714 — Tier 2 (institutional press; MIT reporting its own peer-reviewed study; the fetched, quoted source for the published percentages)
- Shakked Noy and Whitney Zhang · “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence” (working paper, not peer reviewed) · MIT Department of Economics · March 2, 2023 · https://economics.mit.edu/sites/default/files/inline-files/Noy_Zhang_1.pdf — Tier 1 (primary research artifact; fetched, reports the same result in standard-deviation units on 444 participants)
ChatGPT (OpenAI) provided to the treatment group for occupation-specific writing tasks; output quality graded by independent evaluators in the same occupations; assignment was randomized and the design preregistered at the AEA RCT Registry
- Status
- pending
- Method
- Every load-bearing figure is quoted verbatim from MIT's institutional news release about the peer-reviewed paper, 'Study finds ChatGPT boosts worker productivity for some writing tasks' (MIT News, July 14, 2023), fetched this session and saved to sources/mit-news-chatgpt-productivity.html, then grep-verified for '453 college-educated', 'by 40 percent' and 'quality rose by 18 percent'. The peer-reviewed article of record is Shakked Noy and Whitney Zhang, 'Experimental evidence on the productivity effects of generative artificial intelligence,' Science 381:187-192 (July 2023, DOI 10.1126/science.adh2586); its article page and the SSRN and Ovid mirrors returned bot-blocks (HTTP 403/402) to fetchers this session, so the published percentages are quoted from MIT's own release rather than the article body. The authors' pre-peer-review working paper (MIT Department of Economics, March 2, 2023) was fetched and saved to sources/noy-zhang-mit-econ-working-paper.pdf, and it independently reports the same result in standard-deviation units (time down 0.8 SD, quality up 0.4 SD) on a slightly smaller pre-revision sample of 444 participants. Wayback archiving was unreachable from the research box this session, so snapshots are the local sources/ copies and the archive gap is flagged for the checker. No confirmation was sought from OpenAI or the researchers; only the public record was used. Date of verification: September 5, 2026.
- Provider
- ChatGPT (OpenAI), used as a writing assistant; the study was designed and run by an independent MIT research team (Shakked Noy and Whitney Zhang) and published in Science
- Client
- 453 college-educated professionals (marketers, grant writers, consultants, data analysts, HR professionals and managers), studied by MIT and published in Science (July 2023) · cross-industry
- Disclosure
- named
What did the MIT ChatGPT productivity study find?
In a preregistered randomized controlled trial of 453 college-educated professionals, published in Science in July 2023, giving workers access to ChatGPT for occupation-specific writing tasks decreased the time it took to complete them by 40 percent and raised output quality, as graded by independent evaluators, by 18 percent.
Did ChatGPT help everyone equally?
No. The study found that performance inequality between workers decreased: workers who received a lower grade on the first task (done without ChatGPT) benefited more from ChatGPT on the second task, so the tool compressed the productivity distribution by helping lower performers most.
How rigorous is this evidence?
It is a preregistered, incentivized randomized controlled trial published in the peer-reviewed journal Science. Its main caveat is scope: it measured short (20-to-30-minute) writing tasks in a controlled online setting in early 2023, not sustained real-world job performance, and it studied ChatGPT specifically, not every generative-AI tool.