Engineering

The AI Coding Study Nobody Wanted

How a randomized trial found experienced developers got 19% slower with AI tools — and still believed the tools helped.

A craftsman at a clockwork desk surrounded by tangled mechanical helper-arms
Illustration: Edgewisely

A controlled study found AI coding assistants slowed experienced developers down. They didn't notice.

This is an analysis piece grounded in a peer-reviewed randomized study; the interpretation and business implications are Edgewisely's own.

Sixteen experienced open-source developers agreed to be timed. Half worked on real coding tasks with AI assistants enabled, half without. When the researchers at METR, an AI evaluations nonprofit, tallied the results, the group using AI tools finished their tasks 19% slower than the group that didn't, according to the published study. That alone is a striking result. What happened next is the more useful one for anyone deciding how to deploy these tools at a company: after finishing, the developers who'd used AI estimated it had made them roughly 20% faster.

The gap between feeling and measuring

Before the study started, participants predicted AI would speed them up by about 24%. After finishing — measurably slower — they still believed AI had helped by roughly 20%. That's not a small miscalibration; it's a near-total inversion of reality that persisted even after direct, first-hand experience with the outcome. Screen-recording data from the study found part of the explanation: developers spent about 9% of total task time reviewing and correcting AI-generated code, and that review overhead, combined with time spent prompting and waiting for generations, ate up any time genuinely saved on typing and initial debugging.

Why the effect concentrates on senior engineers

The study's participants were experienced developers working in large, mature, real-world codebases with existing architecture constraints, legacy dependencies, and production reliability requirements — exactly the conditions where AI code generation has the least context to work with. A model trained broadly across public code doesn't know a specific company's fifteen-year-old billing system, its undocumented edge cases, or the reason a particular abstraction was chosen over an obvious alternative. Junior developers, by contrast, tend to benefit more from assistant-style tools on isolated, well-specified tasks where that missing context matters less.

For engineering leaders

This is not an argument against AI coding tools. It's an argument against measuring their value by adoption rate or developer sentiment instead of actual task completion time on real work. A team that has rolled out Copilot or Cursor company-wide and reports high satisfaction scores may still be shipping software more slowly than before, if the tools are being used on exactly the kind of large-codebase, high-context work where the METR study found the slowdown concentrated. The fix isn't banning the tools — it's being precise about where they help: unfamiliar codebases, boilerplate, test scaffolding, and onboarding, where the lack-of-context problem cuts the other way and AI genuinely accelerates ramp-up.

For AI coding-tool vendors

The commercial incentive here runs directly against the finding. Vendors are paid on seats and usage, not on verified task-completion speed, and self-reported productivity surveys — the kind vendors love to cite — are exactly the measure the METR study shows is unreliable. Vendors serious about proving value to skeptical enterprise buyers have an opening: publish task-completion-time studies on realistic, large-codebase work, not synthetic benchmarks or user-satisfaction scores, because right now the most rigorous public study available found the opposite of what most marketing claims.

Why developers trust the tools anyway

Stack Overflow's own developer survey data shows a meaningful share of developers say they actively distrust the accuracy of AI-generated code, yet usage keeps climbing regardless. That combination — distrust of output alongside heavy reliance on the tool — is consistent with what the METR researchers observed: the feeling of offloading cognitive work (typing, syntax recall, boilerplate) registers as speed, even when the actual time spent reviewing, correcting, and re-prompting cancels it out. It's the same illusion that makes a long highway drive with cruise control feel faster than a shorter drive spent constantly braking, even when the total time is identical.

Edgewisely has covered the adjacent labor-market question in why AI is quietly erasing the entry-level job, and the tooling landscape itself in the 7 best AI coding assistants of 2026.

The takeaway

If your engineering organization is measuring AI coding-tool ROI by survey, seat count, or gut feel, you likely have no reliable read on whether the tools are actually making your team faster on the work that matters most: maintaining and extending complex, production systems. The METR study is one experiment, on sixteen developers, in one slice of software work — it shouldn't be treated as gospel for every team and every codebase. But it's a rigorous, randomized result that directly contradicts the near-universal assumption that AI coding assistants are an unambiguous productivity win, and it deserves more weight in procurement decisions than another vendor case study.

The tools that feel fastest aren't always the tools that finish fastest. Measure the clock, not the feeling.

Frequently Asked Questions

What did the METR study actually find about AI coding assistants?

METR ran a randomized controlled trial with experienced open-source developers completing real coding tasks, some with AI tools enabled and some without. The group using AI finished 19% slower on average, even though before the study they predicted AI would make them about 24% faster, and after finishing they still estimated it had made them roughly 20% faster.

Why did AI tools slow down experienced developers instead of speeding them up?

Screen-recording data showed developers spent roughly 9% of total task time reviewing and correcting AI-generated code. Combined with time spent writing prompts and waiting for outputs, this review overhead offset any time saved on typing or initial debugging, particularly in large, mature codebases where the AI lacked context about existing architecture and legacy constraints.

Does this mean AI coding tools are not useful at all?

No. The slowdown was concentrated among experienced developers working in complex, mature codebases with strict architectural and reliability requirements. Junior developers and tasks involving unfamiliar code, boilerplate, or onboarding tend to benefit more, since AI's lack of company-specific context matters less on those kinds of tasks.

What should engineering leaders do differently based on this research?

Engineering leaders should measure AI coding-tool value by actual task-completion time on representative work rather than adoption rates or developer satisfaction surveys, since the study found a sharp disconnect between how fast developers felt and how fast they actually were. Rollouts should be targeted at the use cases most likely to help, like onboarding and boilerplate, rather than assumed to help uniformly across all engineering work.

Editor's note — sources: METR ("Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," arXiv:2507.09089), Stack Overflow developer survey data on AI code trust.

Get Edgewisely in your inbox

Business stories that matter, free. Enter your email — no password, no account to set up.
jamie@example.com
Subscribe