Zak FentonWorking notes

· Note

Faster, they said. Slower, they were.

Why I'd check a team's estimate of the time AI saves them

In early 2025, METR ran a randomised trial with 16 experienced open-source developers working on their own repositories: 246 real tasks, using the AI tools of the time. Before doing the tasks, the developers forecast that AI would reduce completion time by 24 per cent. Afterwards, they estimated a 20 per cent reduction. Recorded task times showed a 19 per cent increase.1

What happened next

METR reported a follow-up on 24 February 2026, from an experiment begun in August 2025. Returning developers were estimated to take 18 per cent less time with AI, with an interval from 38 per cent less to 9 per cent more. For new recruits, the estimate was 4 per cent less, with an interval from 15 per cent less to 9 per cent more. Both intervals include no change.2

METR says selection and timing problems make these results unreliable as an estimate of the current effect. Developers were leaving out tasks they did not want to do without AI, and parallel agent work made time harder to record. METR thinks gains probably increased, but these data cannot establish their size.

The bit worth keeping

The group’s average estimate of time saved pointed in the opposite direction from its recorded task times. That’s why I’d check the time a task takes, as well as asking how useful the tool feels. I’ve asked that second question in meetings for years. It produces a warm feeling and nothing you could put in a spreadsheet.

What I’d do with this

In this study, the developers’ estimates did not match the measured task times. I’d check task time and quality before treating a team’s estimate of time saved as an outcome.

Count the checking. Reviewing what a tool produced takes time, and it’s easy to leave it out of the sum.

Treat enjoyment as its own outcome. Worth having, and worth measuring. It’s a poor stand-in for output. I enjoy a lot of things that produce nothing.

Footnotes

  1. Becker, J., Rush, N., Barnes, E. and Rein, D. (2025). Measuring the impact of early-2025 AI on experienced open-source developer productivity. METR, 10 July 2025, report; preprint arXiv:2507.09089. Not peer reviewed.

  2. METR (2026). We are changing our developer productivity experiment design, 24 February 2026, post.