← All case files
pending deployment consulting · US · cross

The 'jagged frontier' experiment: GPT-4 let 758 BCG consultants finish 12.2% more tasks 25.1% faster, yet made them 19% less likely to be right on a task outside AI's reach

In a pre-registered field experiment published in Organization Science, 758 Boston Consulting Group consultants using GPT-4 completed 12.2% more tasks 25.1% faster with higher quality on 18 tasks inside AI's 'frontier', but on one complex task chosen to sit outside it, consultants using AI were 19% less likely to reach the correct answer.

MetricBeforeAfter
Tasks completed and speed on 18 within-frontier tasks, GPT-4 vs no AI No-AI control group 12.2% more tasks completed, 25.1% faster, significantly higher quality
Likelihood of a correct solution on one task outside the frontier, GPT-4 vs no AI No-AI control group 19% less likely to be correct with AI

Verification status: CHECKING — handed to the checker, not yet graduated. NOT verified, NOT green. This is an independent research finding drawn from a pre-registered field experiment and its published preprint; no green badge is claimed and none is implied.

The problem

As GPT-4 arrived inside knowledge-work firms, the debate over its effect on skilled professionals ran on impressions and demos rather than on measured, randomized evidence (source). To test the effect directly, researchers from Harvard Business School, Wharton, MIT Sloan, Warwick and Boston Consulting Group ran a pre-registered field experiment on realistic consulting work (source). The Harvard Crimson reported that the study set “758 BCG consultants” against “18 realistic consulting tasks” (source).

What was built

The paper describes a controlled comparison rather than a deployment: “The preregistered experiment involved 758 knowledge workers. After establishing a performance baseline on similar tasks, subjects were randomly assigned to one of three conditions: no AI access, GPT-4 AI access, or GPT-4 AI access with a prompt engineering overview” (source). The design split the work into tasks inside and outside what the authors call a “jagged technology frontier” to “describe the uneven impact of artificial intelligence (AI) capabilities, where AI assistance improves performance for some tasks but worsens it for others” (source).

The outcome

Inside the frontier, the effect was large and positive: on “18 realistic knowledge tasks within the frontier of AI capabilities”, subjects using AI were “completing 12.2% more tasks and completing them 25.1% more quickly on average while also delivering solutions of significantly improved quality” (source). Independent coverage reported the same headline: consultants “who used GPT-4 completed on average 12.2 percent more tasks, 25.1 percent quicker” (source). A second independent outlet, VentureBeat, reported the same in-frontier result, noting consultants “completed 12.2 percent more tasks on average, and completed tasks 25 percent more quickly” (rounding the study’s 25.1% speed figure to 25%) (source), and confirmed the sample: “The study included 758 consultants, or 7 percent of the consultants at the company” (source).

The counterpart finding, and the reason the study matters, sits outside the frontier: “for a complex managerial task selected to be outside the frontier, subjects using AI were 19% less likely to produce correct solutions compared with those without AI, pointing to potential limitations of AI supporting knowledge workers” (source). The Crimson stated it plainly: “consultants using AI for tasks considered outside of the frontier were 19 percent less likely to produce the correct solutions” (source). A second independent outlet, Forbes, reported the same 19-point effect in the inverse framing, noting consultants “were 19 percentage points more likely to produce incorrect solutions” outside the frontier (source).

What the finding does and does not show

The quality gains were not spread evenly, but the precise size of the quality effect is where the popular retelling drifts from the record. The peer-reviewed abstract states only “significantly improved quality” and puts no single percentage on it (source). The often-repeated “40% higher quality” line is not in the abstract; in a recap, study co-author Karim Lakhani describes the skill skew as “quality scores rising 43% compared to 17% for the highest-skilled participants” (source). That last figure is the weakest load-bearing source here: it is a Tier 3 first-party blog by one of the authors, not the peer-reviewed text, so this page treats the quality result qualitatively as the published abstract states it and flags the by-skill split as an author’s recap, not a verified headline number. Read precisely, the verified claims are the three the peer-reviewed publication states directly: +12.2% tasks, 25.1% faster inside the frontier, and 19% less likely to be correct outside it.

How this was verified

Method: every figure was retrieved live on 2026-08-30 from the peer-reviewed Organization Science (INFORMS) article (DOI 10.1287/orsc.2025.21838, published online 2026-03-11) and cross-checked against The Harvard Crimson’s independent 2023 report and the SSRN working-paper preprint. Each quotation is verbatim; the published article and the preprint were captured on the Wayback Machine and saved locally in sources/. The critical figures are stated by an independent peer-reviewed publication, not by the AI vendor or by BCG, which is the strongest available form of independence for this claim.

Sources

  1. Dell’Acqua, McFowland, Mollick, Lifshitz-Assaf, Kellogg, Rajendran, Krayer, Candelon, Lakhani · “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality” · Organization Science (INFORMS), Vol. 37 No. 2, published online 2026-03-11 · Tier 1 (primary, independent peer-reviewed publication) · https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838
  2. The Harvard Crimson (Martinez, Mezitis) · “Harvard Business School Partners with BCG on AI Productivity Study” · 2023-10-13 · Tier 2 (independent press reporting the study) · https://www.thecrimson.com/article/2023/10/13/jagged-edge-ai-bcg/
  3. Dell’Acqua et al. · “Navigating the Jagged Technological Frontier” (SSRN working paper 24-013, preprint of the same study) · 2023 · Tier 1 (primary preprint) · https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321
  4. Karim Lakhani · “Discovering AI’s jagged frontier — and what we’ve learned since” (author’s recap) · 2026-03-16 · Tier 3 (first-party author blog; used only for the by-skill quality nuance) · https://professorkl.substack.com/p/discovering-ais-jagged-frontier-and
  5. VentureBeat (Matt Marshall) · “Enterprise workers gain 40 percent performance boost from GPT-4, Harvard study finds” · 2023-09-25 · Tier 2 (second independent press outlet; corroborates the 12.2% in-frontier figure and the 758-consultant sample, independent of OpenAI, BCG and the study authors) · https://venturebeat.com/ai/enterprise-workers-gain-40-percent-performance-boost-from-gpt-4-harvard-study-finds
  6. Forbes (Dan Pontefract) · “Harvard And BCG Unveil The Double-Edged Sword Of AI In The Workplace” · 2023-09-29 · Tier 2 (second independent press outlet for the out-of-frontier 19 figure; states the 12.2%/25.1% in-frontier figures too, independent of OpenAI, BCG and the study authors) · https://www.forbes.com/sites/danpontefract/2023/09/29/harvard-and-bcg-unveil-the-double-edged-sword-of-ai-in-the-workplace/

GPT-4 (OpenAI); one arm also received a prompt-engineering overview

Verification record
Status
pending
Method
Retrieved live this session from the peer-reviewed Organization Science (INFORMS) article (DOI 10.1287/orsc.2025.21838) and its SSRN working-paper preprint (24-013), cross-checked against two independent press outlets (The Harvard Crimson and VentureBeat). Every figure quoted verbatim and archived to the Wayback Machine, with local copies in sources/. The headline figures are stated by an independent peer-reviewed publication, not by the AI vendor or by BCG's marketing.
Provider
GPT-4 (OpenAI), studied in a field experiment run by researchers from Harvard Business School, Wharton, MIT Sloan, Warwick and BCG
Client
758 Boston Consulting Group management consultants (the 'Navigating the Jagged Technological Frontier' field-experiment cohort) · consulting
Disclosure
named
Questions this file answers
How much more productive were BCG consultants using GPT-4 in the study?

On 18 realistic tasks inside the 'frontier' of AI capability, consultants using GPT-4 completed 12.2% more tasks and completed them 25.1% more quickly on average than consultants without AI, with significantly improved quality.

What is the 'jagged frontier' finding?

AI helps unevenly. On a complex managerial task deliberately chosen to sit outside AI's frontier, consultants using AI were 19% less likely to produce a correct solution than those working without it, so the same tool that lifted in-frontier work degraded out-of-frontier work.