METR’s early-2025 experiment found that experienced open-source developers completed their assigned tasks 19 percent more slowly with AI access. Its February 2026 follow-up is a warning against carrying that finding forward as a permanent verdict. The researchers said the new study could not reliably establish the tools’ current productivity effect.[1][2]
Some developers declined to participate because they did not want to work without AI. Reduced compensation introduced another possible selection effect, while developers running several agents made time measurement harder. The remaining participants were no longer a clean window into everyone using the tools.[2]
The raw follow-up estimates pointed toward faster completion, but their confidence intervals included no improvement. METR said conversations with participants suggested greater benefits than before, while stressing that the data offered only weak evidence about the size of that change.[2]
Productivity has more than one denominator
An earlier workplace study supplies a different result, and a different setting. In their revised 2023 working paper, Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied the rollout of an assistant to 5,179 customer-support agents. Issues resolved per hour rose 14 percent on average. Gains were much larger for novice and lower-skilled workers, while experienced, highly skilled agents saw little change. This was assistance inside a repeatable customer-service workflow, not autonomous software development.[4]
The two findings need not contradict each other. An expert working in a familiar repository may spend time correcting suggestions that do not fit its conventions. A newcomer facing a recurring support problem may benefit from ready access to patterns used by experienced colleagues. That is a plausible interpretation of the contrast, not a mechanism proved by comparing two unrelated studies. The worker, task, tool and measure changed together.
Anthropic’s January 2026 experiment asked a third question: what happens to learning? Researchers recruited 52 mostly junior engineers to use an unfamiliar Python library. The AI group averaged 50 percent on a subsequent comprehension quiz, compared with 67 percent for the group coding by hand. The modest completion-time advantage was not statistically significant. The experiment measured immediate understanding in a short exercise; it did not establish that every use of an assistant causes lasting skill loss.[3]
Its qualitative analysis also found different patterns of use. Asking conceptual questions and seeking explanations was associated with stronger comprehension than delegating the work wholesale. Those behavioral groups were small and were not separately randomized. They suggest something worth testing, rather than a guaranteed recipe for retaining expertise.[3]
The work continues after the answer
For a team choosing a tool, the useful unit is an accepted result. A patch that appears quickly but takes an hour to understand has not saved an hour. Conversely, a prototype that makes a previously unaffordable investigation possible can create value even if it does not accelerate a standard ticket. A sensible evaluation should make room for both effects without confusing them.
One practical approach would separate elapsed time from human attention. Record the interval to a reviewed result, the minutes spent supervising it, the corrections required, and whether another person can maintain the output. Include a small delayed check of understanding when the task introduces unfamiliar concepts. These are proposed evaluation measures, not results from a new Daybreak benchmark.
That distinction becomes especially useful when several agents run at once. Lower human effort could coexist with longer elapsed time; more output could coexist with more review work. A single speed percentage collapses decisions that a developer or manager may care about separately.
The enduring lesson of this research is to attach every productivity claim to its conditions. Which workers? Which tasks? Which model and interface? What counted as finished? The evidence can support real gains and real costs at the same time. It becomes less useful when a result from one setting is promoted into a permanent verdict on an entire technology.
Sources & further reading
Original reporting and research behind this article.
- METR: randomized developer studyJul 10, 2025
- METR: changes to its productivity experimentFeb 24, 2026
- Anthropic: AI assistance and coding skill formationJan 29, 2026
- Brynjolfsson, Li and Raymond: Generative AI at Work, revised working paperWorking paper revised November 2023