Leading Through AI Adoption · Session 19
The AI Perception Gap
A randomized trial found experienced developers were 19% slower using AI tools, while believing they had been 20% faster. When self-report and reality diverge that badly, most of what you hear about AI's impact is unreliable.
In 2025, the research group METR ran a randomized controlled trial on experienced open-source developers. Not students, not a lab exercise. Sixteen developers working 246 real issues in large repositories they already maintained, projects averaging over 22,000 stars and a million lines of code.
Before starting, the developers predicted AI tools would make them about 24% faster.
They were 19% slower.
Afterwards, asked to estimate what had happened, they said AI had made them roughly 20% faster. They had just personally lived through the opposite and could not perceive it.
Worth being honest about the limits, because they're substantial. Sixteen developers is a small sample. They were experts working in codebases they knew deeply, which METR explicitly say does not represent most software work, and they were using early-2025 tooling with limited prior experience of it. The authors state plainly that their results should not be generalized, and they revised the experiment design in early 2026. Do not walk into a meeting and announce that AI makes everyone slower.
But the gap between prediction, reality, and recollection is the finding that matters for leaders, and that gap is not really about AI at all.
Why the gap exists
Nothing here requires anyone to be dishonest or unusually foolish. The mechanism is ordinary.
AI removes the friction you notice and adds the friction you don't. Staring at an empty file is memorable effort. It feels like work, and it's the part AI eliminates. Reading generated code, checking whether it does what it claims, discovering it subtly doesn't, and correcting it is diffuse effort. It doesn't register as a distinct event you'd recall later. So the felt experience is "that was easier," even when the clock says it took longer.
Effort and time get confused. People routinely report the less effortful path as the faster one, because that's what the experience felt like. Those aren't the same measurement.
Nobody has a control group for their own week. You never observe the version of yourself who did the same task without the tool. There's nothing to compare against except a guess, and the guess is shaped by what everyone around you is currently saying.
This is the same territory as the Dunning-Kruger research, where the more durable finding isn't that incompetent people are uniquely overconfident but that self-assessment is unreliable across the board, including from your strongest people. Expertise doesn't repair it. The METR participants were experts and were confidently wrong about their own week.
What this breaks in how you lead
Most engineering organizations run almost entirely on self-report. Standups, retros, one-on-ones, status updates, the whole apparatus is people telling you how it's going. That has always been imperfect. AI has widened the error bar considerably on one specific class of question.
Three practical consequences:
Enthusiasm is not evidence. When a team reports AI is transforming their productivity, they are reporting a genuine feeling. They are not reporting a measurement, and the METR data suggests the feeling can point the wrong direction entirely. Treat it as a signal worth investigating, not a result.
Skepticism isn't evidence either. The reverse applies. An engineer certain AI is useless is also self-reporting. Their conviction deserves the same scrutiny as the enthusiast's, no more and no less.
Executive confidence outruns everyone's. Surveys consistently find a large trust gap between executives and the people doing the work. The further you sit from the keyboard, the more your impression is built from other people's impressions, which were themselves unreliable. Confidence compounds as it travels up.
Measuring instead of asking
You don't need a research lab. You need to stop treating "how's it going" as data.
- Watch outcomes, not sentiment. Lead time and change failure rate don't have feelings about AI. If something genuinely improved, it eventually shows up there.
- Compare like with like, over time. Not team A against team B, whose work differs. The same team, same class of work, before and after.
- Ask about specifics, not impressions. "Did AI help this week" gets you a vibe. "Walk me through the last thing it saved you real time on, and the last thing it cost you time on" gets you two concrete events you can actually reason about.
- Ask what it cost. Almost nobody volunteers the debugging session chasing a plausible-looking hallucination, because it doesn't feel like a distinct event. It has to be asked for directly.
The uncomfortable part
This applies to you too.
If you have a strong conviction about whether AI is helping your organization, and that conviction rests mainly on how things feel and on what people tell you in meetings, you're relying on the exact instrument this study found unreliable. Being the leader doesn't exempt you. It usually means you're further from the work and more dependent on secondhand impressions than anyone else in the building.
For Discussion
Write down your current belief about whether AI is helping your team, and then write down what specifically that belief is based on. If the honest answer is "impressions and what people have told me," what measurement would actually settle it?
Sources and Further Reading
- Thinking, Fast and Slow by Daniel Kahneman
Why self-assessment is unreliable in general, which is the mechanism underneath the METR result in this session.
These are Amazon affiliate links. If you buy through one, I earn a small commission at no extra cost to you. I only link books I actually use and recommend.