IAMUVIN

Industry Analysis & Trends

Does AI Make Developers Faster? What the Measurements Say

Uvin Vindula·August 14, 2026·16 min read
Share

TL;DR

We do not have a clean measurement of whether AI makes experienced developers faster, and the 2026 evidence says so more clearly than the 2025 evidence did. METR's original randomised trial, published 10 July 2025, put 16 experienced open-source developers across 246 tasks in repositories they had worked in for around five years, and measured them 19% slower in the AI-allowed condition, with a confidence interval from 2% to 39% slower. Those developers had forecast a 24% speedup and reported a 20% speedup afterwards.

The follow-up, published 24 February 2026 with 57 developers, 143 repositories and over 800 tasks, is the number nobody quotes. The 10 developers who were also in the original study came out 18% slower, interval from 38% slower to 9% faster. The 47 newly recruited developers came out 4% slower, interval from 15% slower to 9% faster — statistically indistinguishable from zero. METR is changing the experiment design because selection effects make the data, in its own words, "only very weak evidence for the size of this increase" — METR's stated position is that developers are likely more sped up in early 2026 than in early 2025, and that these measured numbers are a lower bound.

Meanwhile METR's May 2026 survey of 349 technical workers found a median self-reported 1.4x to 2x gain in the value of their work, and a median 3x on speed — the speed figure METR itself expects to be the inflated one. The gap between that and a measured effect near zero is the finding.


The 19% Slowdown Study Has a 2026 Sequel

If you have seen one number in this argument, it is minus 19%. It comes from METR's randomised controlled trial published 10 July 2025:

  • 16 experienced open-source developers.
  • 246 tasks.
  • Repositories in which participants had roughly five years of prior experience.
  • Measured result: 19% slower in the AI-allowed condition, confidence interval from 2% to 39% slower.
  • Forecast before the study: 24% faster.
  • Self-report after the tasks: 20% faster.

That triple — forecast 24% faster, felt 20% faster, measured 19% slower — is why the study travelled. It is a roughly 40-percentage-point gap between perception and measurement, in a population that should be hard to fool.

Then almost nobody read what happened next. METR published an uplift update on 24 February 2026 with a much larger sample:

CohortParticipantsMeasured effectConfidence interval
Also in the original study1018% slower38% slower to 9% faster
Newly recruited474% slower15% slower to 9% faster

Total: 57 developers, 143 repositories, over 800 tasks.

Read the second row carefully. A 4% slowdown with an interval spanning 15% slower to 9% faster is not a slowdown finding. It is an absence of a finding. The data is consistent with a modest slowdown, with no effect, and with a modest speedup.

The first row keeps the original effect size — 18% slower for the returning developers — but the interval now crosses zero too. On its own, that row cannot carry the claim the 2025 paper carried.

So the honest 2026 position is neither "AI makes experienced developers 19% slower" nor "AI makes developers 3x faster". It is that we currently have no clean measurement. The study most people cite as proof of the first claim has a larger successor whose intervals cross zero, and whose authors say the selection effects in it push the estimate downwards — so their reading is that the true figure sits above the measured one, not below it.

Why METR changed the design

This part is more useful to a practitioner than either effect size, because it tells you why measuring this is hard in your own team.

METR states three reasons the data amounted to only very weak evidence for the size of the speedup it believes is there:

  1. Developers refused to enrol. METR saw "a significant increase in developers choosing not to participate in the study because they do not wish to work without AI", and 30% to 50% of those who did enrol said they withheld tasks they did not want to attempt unaided. Both effects strip out the work with the highest expected AI uplift, which is why METR calls its estimate a lower bound.
  2. The pay rate dropped. Compensation fell from $150/hr in the original study to $50/hr in the follow-up, which METR says likely contributed to the same selection.
  3. Agentic tools broke self-reported time. Developers worked on other things while waiting for an agent, so METR's "measurements of time-spent on each task are unreliable for the fraction of developers who use multiple AI agents concurrently". Self-reported time-spent stops meaning wall-clock time and stops meaning attention either.

The second one kills the most common internal measurement approach on the spot. If your team tracks "time on task" through self-report or through an IDE timer, agentic workflows have already made that metric unreliable. Waiting is no longer idle, so a task that takes 90 minutes of wall clock and 20 minutes of attention gets recorded as whichever number the developer happens to think of.

What Developers Say Against What Was Measured

METR published an AI usage survey on 11 May 2026 covering 349 technical workers, including 87 software engineers, 71 researchers, 129 academics and PhD students and 48 founders and managers.

  • Median reported change in speed: 3x.
  • Median reported change in the value of their work: 1.4x to 2x.
  • Retrospective self-estimates on value: 1.3x in March 2025, 2x at survey time, forecast 2.5x by March 2027.
  • Higher perceived gains among heavier AI users and among Claude Code users.
  • METR's own employees reported notably lower gains. METR lists four possible explanations for that, including the opposite readings that its staff are better calibrated on the earlier perception-measurement gap and that they overindex on it, and says its intuition points weakly at the first.

And METR's caveat in the same post, stated plainly: "survey results are not necessarily grounded in reality", citing the 2025 study's 40-percentage-point gap between perceived and actual.

That last bullet is the one I find hardest to argue with. The population most aware that perceived speedup diverged from measured speedup reported smaller gains than everyone else. That is either calibration or self-consciousness, and there is no way to tell which from a survey.

Why Both Numbers Can Be Honest

A 3x self-report and a measured effect near zero look like someone is lying. Neither has to be.

Speed and value are different quantities and the survey measured both separately: median 3x on speed, median 1.4x to 2x on value. If a tool triples the rate at which you produce work whose value roughly doubles, and some of that work is work you would not have attempted at all, a task-completion RCT on pre-selected tasks will not see most of it. The RCT measures time to finish a chosen task in a repository you know well. That is a narrow slice.

Three mechanisms plausibly explain a felt speedup that a controlled task measurement does not capture:

  • Task substitution. You attempt things you would previously have skipped or delegated. The counterfactual is not a slower version of the same task; it is no task.
  • Effort substitution. The work feels easier even at equal duration. Reduced cognitive load reads as speed and does not appear in elapsed time.
  • Waiting reclassified as working. As METR found, agentic tools turn blocked time into parallel time, which makes an unchanged wall-clock number feel shorter.

None of that makes the 3x figure a measurement. It makes it a perception with identifiable causes, which is a more useful thing than either a dismissal or a headline.

Why Measured Gains Lag Benchmark Gains

The benchmark side of this looks nothing like the field side, and two 2026 papers explain the gap.

The Stanford AI Index 2026 records SWE-bench Verified climbing from 60% to nearly 100% of human baseline in a single year, Humanity's Last Exam gaining 30 percentage points, GPQA passing an 81.2% human expert baseline to reach 93%, and OSWorld agent task success rising from 12% to around 66%. The same report holds up a counter-example from the same model class: reading analog clocks at 50.1% accuracy.

What blocks production is not the benchmark axis:

  • CL-Bench (arXiv 2606.05661, June 2026) tested continual learning across six expert-validated domains — software engineering, signal processing, disease outbreak forecasting, database querying, strategic game-playing and demand forecasting — where tasks share a learnable latent structure a stateful system can discover online. The best system reached only a 25.4% normalised gain over its stateless baseline, and naive in-context learning beat systems dedicated to memory management on most tasks. I would rate this medium confidence as a single benchmark paper, but the direction matches what I see building agents.
  • ["The Long-Horizon Task Mirage?"](https://arxiv.org/html/2604.11978v1) (Wang, Bai, Sun and colleagues, arXiv 2604.11978v1, 13 April 2026) analysed over 3,100 evaluated trajectories across Web, OS, Embodied and Database domains using GPT-5 variants and Claude-4-Sonnet. It attributed 72.5% of failures to process-level causes and 27.5% to design-level ones, found performance dropping sharply past a threshold rather than degrading linearly, and identified planning decomposition failures and catastrophic forgetting as dominant. The authors conclude that scaling base models alone is unlikely to resolve the dominant failure mechanisms.

There is also a professional-task measurement that sits between benchmarks and payroll data. GDPval, an OpenAI benchmark released in September 2025 and published at ICLR 2026, covers 1,320 tasks across 44 occupations in the top nine US-GDP industries, each built from real work products and vetted by professionals averaging 14 years of experience, then scored by blind pairwise expert comparison. In the original results Claude Opus 4.1 reached a 47.6% win-or-tie rate against human experts, GPT-5-high 38.8% and o3-high 34.1%. OpenAI's accompanying claim that models complete these tasks roughly 100x faster and 100x cheaper reflects pure inference time and API billing rates, and excludes human oversight, iteration and integration.

That exclusion is the whole gap between a benchmark result and a delivered feature. A 47.6% win-or-tie rate meant the expert preferred or tied the model's output on fewer than half the tasks — but that was a September 2025 snapshot and it has not held: on the same gold subset GPT-5.2 Thinking now reaches 70.9% wins or ties, Claude Opus 4.5 59.6% and Gemini 3 Pro 53.5%. What did not move is the cost of checking. The graders' complaints were rarely factual errors, which run at 2% to 3.5% across models; they were instruction adherence and formatting, which is exactly the class of defect a human has to catch before the work ships. The 100x figure counts the part that got cheap. The paper prices the part that did not: under a "try the model, and if it is still unsatisfactory fix it yourself" workflow, GPT-5 came out 1.39x faster and 1.63x cheaper than the unaided expert, against naive ratios of 90x and 474x. Two orders of magnitude of the headline is review time.

That is the mechanical reason a developer can watch benchmark charts go vertical and still not measure a speedup on their own repository. The thing improving fastest and the thing blocking delivery are not the same thing. Most of the recoverable gain sits in the harness rather than the model, which is the argument I make in why your agent benchmark number is mostly your harness and in context engineering for production.

The Jobs Data: Two Datasets, One Composition Change

Employment is where this debate gets sloppiest. One camp cites Stanford and declares entry-level software work finished. The other cites Indeed and declares the panic over. Both datasets are credible, both were published in 2026, and together they describe a composition change rather than a headcount collapse.

Stanford: the young-worker gap widened from 15% to 19%

The revised "Canaries in the Coal Mine?" paper from the Stanford Digital Economy Lab, published 12 August 2026, uses ADP payroll data from November 2022 to June 2026. Its six findings:

  1. No widespread economy-wide displacement.
  2. Workers aged 22 to 25 in AI-exposed occupations sit around 19% below where they would be had they kept pace with less-exposed peers.
  3. That gap widened from 15% in July 2025 to 19% in June 2026.
  4. The adjustment runs through reduced hiring, not separations.
  5. Declines concentrate where AI automates rather than complements; complementary occupations are flat or rising.
  6. Employment adjusts before base pay.

In absolute terms, employment of 22 to 25 year olds in the two most exposed quintiles fell about 11% between November 2022 and June 2026. An earlier release reported that 22 to 25 software developer employment in ADP was down nearly 20% from its late-2022 peak to July 2025.

Point 4 is the one that changes what you do with this. Reduced hiring is not the same labour-market event as layoffs. It is slower, quieter, and it lands entirely on people who are not yet employed — which is why it shows up in a payroll panel before it shows up in a news cycle.

Point 5 is the one a developer can act on, with one qualifier attached. The decline concentrates where AI automates the occupation, not where it complements it, and complementary occupations are flat or rising — but Stanford measures that rise particularly among more experienced workers, so it is not a job that is currently open at the entry point.

Indeed: postings up 15%, and 71% of the growth is senior

Indeed Hiring Lab, in an 8 July 2026 analysis by Guillermo Gallacher, measures the other side:

MeasureValue
US software development postings since Claude Code launched in late February 2025up nearly 15%
Overall postings over the same perioddown 7%
Software development postings against February 2020about 27.5% below
Overall postings against February 2020essentially unchanged
Share of net increase from May 2025 to May 2026 from senior roles71%
Share of net increase from jobs mentioning AI in the title37%

Software development postings rose by nearly 15% in a period when postings overall fell 7%. That is a real reversal and it is not compatible with a story in which AI is eliminating software work outright. It is also not a recovery: the category is still around 27.5% below February 2020, while the overall market has returned to its February 2020 level.

The composition line is the finding. Of the net increase between May 2025 and May 2026, 71% came from senior roles and 37% from jobs with AI in the title. Growth at the top of the seniority ladder over May 2025 to May 2026, against a 22-to-25 gap that widened from 15% in July 2025 to 19% in June 2026, is one coherent picture, not two contradictory ones: the market is hiring, and it is hiring people who already have the experience.

What Both Authors Say About Causation

Neither team claims AI caused what it measured. That matters more than either headline.

Stanford's stated limitations: the study cannot establish causation; gaps shrink when controlling for education; some patterns predate generative AI; estimates are larger in ADP than in national benchmarks; and results are sensitive to specification. Five caveats, on a paper routinely cited as proof of a causal claim.

Gallacher states plainly that correlation does not imply causation and names other market factors. His 15% rise is measured against a period that begins with Claude Code's late-February 2025 launch. He states that correlation does not imply causation, then calls the timing "a coincidence that cannot be ignored" — a flagged association, not an attribution.

So the strongest supportable claim from both datasets together is this: between November 2022 and mid-2026, entry-level employment in AI-exposed occupations fell relative to less-exposed peers, the gap widened through the period, postings in software development rose while postings overall fell, and the growth concentrated in senior and AI-titled roles. Why any of that happened is not established by either dataset.

Anyone extending those numbers into a forecast is adding the causal step themselves.

The Dated Predictions Worth Checking

One prediction in this area has a hard date and is therefore worth keeping on file. At Davos in January 2026, Dario Amodei said AI models would replace software developers' work "within a year" — so by roughly January 2027 — plus Nobel-level scientific research within two years and 50% of white-collar jobs gone within five (Fortune, 23 January 2026). He put the software claim more loosely elsewhere in the same week, as models doing "most, maybe all" of what software engineers do in six to twelve months, and his written version of the jobs number is narrower: 50% of entry-level white-collar jobs within one to five years.

At the same event Demis Hassabis said current AI is "nowhere near" human-level AGI, put the odds at 50% within the decade, and said "maybe we need one or two more breakthroughs". Yann LeCun said "We're never going to get to human-level intelligence by training LLMs or by training on text only."

Date-stamped, falsifiable predictions are the only kind worth arguing about, and I go through the full set alongside the capability thresholds labs actually govern by in AGI has no definition, capability thresholds do the work.

For calibration on where model capability actually sits, METR's Time Horizon 1.1, published 29 January 2026, measured the task length at which a model succeeds half the time: 320 minutes for Claude Opus 4.5, with an interval of 170 to 729 minutes, against 214 minutes for GPT-5 and 60 minutes for Claude Sonnet 3.7, and a doubling time of roughly 131 days for post-2023 models — a 20% acceleration on the previous 165-day estimate. METR's own caveat on that curve is load-bearing and rarely quoted: only 5 of its 31 long tasks have measured human baselines, the other 26 use estimated times, the confidence intervals "are still very wide", and "the trend in time horizon is somewhat sensitive to task composition."

A 320-minute 50%-reliability horizon and a measured team speedup near zero are not in conflict. A five-hour task completed half the time still needs a person to check it, and the checking is not in the horizon number.

How to Measure This in Your Own Team

Every organisation asking whether AI tools are worth the seat cost is running a worse-designed version of METR's study without knowing it. If you are going to measure, measure the things the 2026 evidence says survive contact with agentic tools.

text
Do not measure:
- self-reported time on task      (agentic waiting breaks it)
- lines of code or PR count       (output volume, not delivered value)
- developer sentiment alone       (the 40-point perception gap)

Do measure:
- cycle time from first commit to merged, per task class
- change failure rate and rollback count
- review time per PR, and review rounds per PR
- tasks attempted that were previously deferred or dropped
- token and inference spend per merged change

Four notes on that list.

Cycle time has to be split by task class. Aggregate cycle time moves when the mix of work moves, and AI tooling changes the mix — that is the substitution effect. If you cannot split it, you are measuring composition, exactly as the employment datasets are.

Change failure rate is the one people skip, and it is where a speedup gets repaid. A faster merge that produces a rollback is negative throughput, and unlike duration it is recorded automatically.

Review load is where agentic output actually lands. More generated code arrives at the same number of reviewers, so review time per PR and review rounds per PR are the first place a real bottleneck shows up.

Spend per merged change is the only figure that converts this to money, and it is the one most teams never build. I break the calculation down in tokens per task, the real AI cost model, and on the integration side in the real cost of AI integration.

Run it for a quarter before deciding anything. METR needed 57 developers, 143 repositories and 800 tasks to produce an interval that still crosses zero. A four-week internal trial on eight engineers will produce a number, and that number will be noise.

What I Tell Junior Developers

The two datasets say something specific about where to stand, and it is neither "the field is closed" nor "nothing has changed".

Hiring is slower at the entry point in AI-exposed occupations — a gap that widened from 15% to 19% between July 2025 and June 2026, running through reduced hiring rather than separations. At the same time, software development postings rose nearly 15% while the overall market fell 7%, with 71% of that net growth in senior roles.

The gap is between those two facts: demand for people who can be trusted with a merge, weak demand for people who cannot yet be. Since the decline concentrates where AI automates rather than complements, the work that holds up is the work that complements — review, integration, debugging across boundaries, and the judgement that decides whether generated code is correct. None of that is entry-level by accident; it is entry-level by convention, and the convention is what moved.

Neither dataset establishes causation, so treat this as a description of the market, not a law of it. I go further into the practical route through a practical roadmap for learning to code in 2026 and the broader context in tech trends from a developer's perspective.

Key Takeaways

  • The famous 19% slowdown has a sequel that does not replicate it cleanly. METR's 24 February 2026 follow-up put 47 newly recruited developers at 4% slower with an interval from 15% slower to 9% faster — indistinguishable from zero.
  • METR changed the design for stated reasons. Selection effects in the AI-disallowed arm and agentic tools making self-reported time-spent unreliable, because developers worked on other things while waiting.
  • The perception gap is the durable result. The 2025 trial found a forecast 24% speedup, a felt 20% speedup and a measured 19% slowdown in the same population.
  • Self-reports remain high and METR distrusts them. Its 11 May 2026 survey of 349 technical workers found a median 3x on speed and 1.4x to 2x on value, with METR stating survey results are not necessarily grounded in reality.
  • The entry-level gap is real and it runs through hiring. Stanford's ADP analysis to June 2026 shows employment of 22 to 25 year olds in AI-exposed occupations around 19% below where it would be had it kept pace with their less-exposed peers, widened from 15% in July 2025, via reduced hiring rather than separations.
  • Postings grew while the gap widened. Indeed measured US software development postings up nearly 15% since late February 2025 against a 7% overall fall, still 27.5% below February 2020, with 71% of the net growth from senior roles.
  • Neither employment author claims causation. Stanford lists five limitations including sensitivity to specification; Gallacher states outright that correlation does not imply causation.

About the Author

I'm Uvin Vindula — a Web3 and AI engineer based between Sri Lanka and the UK. I ship production Next.js, smart contract and AI systems with these tools in the loop every day, which is exactly why I do not trust my own sense of how much faster they make me. You can see my work at iamuvin.com or reach out about a project at hello@iamuvin.com.

If you are trying to work out whether AI tooling is paying for itself on your team rather than in a vendor deck, let's talk about your project.

Working on a Web3 or AI project?

Share

More in Industry Analysis & Trends

All Industry Analysis & Trends articles
Uvin Vindula

Uvin Vindula

Web3 and AI engineer based in Sri Lanka and the UK. Author of The Rise of Bitcoin. Founder of ASI Research Labs. Director of Blockchain and Software Solutions at Terra Labz. Founder of uvin.lk — Sri Lanka's Bitcoin education platform with 10,000+ learners.