Scovai Scovai
AI & Operations 2026-09-18 1 min read

AI Sped Up the Work, Not the Output: MIT Sloan's 500,000-Developer Study Names the Downstream Stage Where Your Productivity Gains Die

DSL

Dr. Sarah Liu

AI Sped Up the Work, Not the Output: MIT Sloan's 500,000-Developer Study Names the Downstream Stage Where Your Productivity Gains Die

Autonomous coding agents raised developer commit activity by a cumulative 240%. Releases โ€” the stage where software reaches a customer โ€” rose 30% (Demirer, Musolff & Yang, NBER, 2026). Same developers, same tools, same study, same window. Eight-tenths of the measured gain never left the building.

That gap is the shape of most AI business cases written this year. The pilot instruments the stage the tool accelerates. Nobody instruments the stage that gates delivery, because that stage has no license fee attached to it.

The paper is called Writing Code vs. Shipping Code, and the distinction in the title is the entire finding. AI productivity gains are real, large, and measured at the task. They arrive at the customer heavily discounted.

What a 500,000-Developer Study Can See That a Pilot Can't

Mert Demirer (MIT Sloan), Leon Musolff and Liyuan Yang matched AI-usage telemetry to more than 500,000 GitHub developers in a matched event-study design, then tracked the effect across successive tool generations (NBER Working Paper 35275, 2026).

The generational results are worth holding separately, because they are the part most readers will want to quote:

  • Autocomplete: +30% cumulative effect on commits
  • Interactive coding agents: +180%
  • Autonomous coding agents: +240%

Each generation is a genuine step change. Anyone arguing that agentic tools don't move task-level throughput is arguing against a very large sample.

Then the attenuation. That 240% commit effect falls to 80% for number of projects and 30% for actual releases. And when the authors moved outside GitHub to four major software marketplaces, they found a sharp increase in the number of new apps and no increase in total usage.

More code written. Somewhat more projects started. A little more shipped. Nothing more used.

The marketplace result deserves separate weight, because it rules out a comfortable reading. A bottleneck story alone would predict a backlog โ€” output waiting behind a gate, recoverable once the gate widens. But new apps rose while total usage did not, which points at something past the release stage: the market absorbed more supply without producing more demand. Some of the missing gain is queued behind human review. Some of it was never going to be value in the first place, because more output of a thing nobody asked for is not output. Both readings should make an operator cautious about a business case that converts task-level speed directly into revenue.

Why AI Productivity Gains Die Before the Last Stage

The authors' explanation is the weak-link hypothesis, and it is older and better-established than anything specific to AI: in a multi-stage production chain, total output is governed by the least-improved stage, not the average one.

Writing code is one stage. Review, integration testing, security sign-off, release approval and deployment are the others โ€” and no generation of coding tool touched them. So the constraint moved. It did not disappear; it relocated to the first human gate downstream of the acceleration.

The Elasticity Number Is the Argument

The paper puts a coefficient on it: an estimated elasticity of substitution of 0.23 between AI and human effort (Demirer et al., NBER, 2026).

Read that plainly. A low elasticity means strong complementarity. AI and the humans downstream of it are not substitutes competing for the same work โ€” they are inputs that need each other in rough proportion. Triple one without touching the other and you do not get triple output. You get a queue.

This is the number to put in front of anyone whose plan quietly assumes that a sufficiently good agent will eventually absorb the reviewer too. The best available estimate says the opposite, and says it with a decimal point.

Your Pipeline Has the Same Shape

Software is the setting, not the scope. The structure that produced this result โ€” sequential stages, machine-acceleratable work at the front, human judgment gates at the back โ€” is the structure of nearly every operating process in a 50โ€“500 FTE company.

Order-to-cash: quote generation is automatable; credit approval and exception handling are not. Hire-to-onboard: sourcing and screening are automatable; the offer decision and the first-week handoffs are not. Ticket-to-resolution: triage and drafting are automatable; escalation judgment is not.

In each case the AI budget sits on the front stage and the capacity ceiling sits on the back one.

How to Find Your Weak Link in an Afternoon

The diagnostic does not require new tooling. Take one process and write down its stages end to end โ€” six or seven is typical. For each stage, mark two things: whether AI touched it in the last twelve months, and what the queue looks like in front of it today.

The weak link is almost always the first stage that answers no to the first question and growing to the second. It will usually be a stage staffed by one or two senior people who were already the escalation point before any of this started, which is precisely why nobody proposed automating it and precisely why it cannot absorb more volume.

Then ask the question that reframes the budget conversation: if this stage processed 20% more units per week, what would that be worth? Compare it to the cost of the next tranche of upstream licenses. In most mid-market pipelines the comparison is not close.

Independent evidence says this misalignment is close to universal. Behavioral telemetry across 120,620 workers found only 2% at workflow-integration maturity โ€” using AI inside a restructured process rather than as a side lookup, with 27% still at simple research assistance (ActivTrak Productivity Lab, 2026). If 98% of adoption is happening around the process rather than inside it, the downstream stages were never in scope to begin with.

Production data points the same way from the opposite end. A study of agent versus search usage found matched tasks completing in 36 minutes against 269 minutes, an 87% time reduction โ€” while the follow-up work shifted upward into verification and extension rather than vanishing (Perplexity & HBS, arXiv, 2026). The human work didn't leave. It moved to the stage you weren't measuring.

The Gate You Didn't Fund Is Also the Expensive One

Two further findings make the downstream stage harder to ignore than a simple queueing problem.

First, it costs more than the tool. McKinsey QuantumBlack's analysis of agentic unit economics found that for a banking customer-service agent, token costs represent just 20โ€“25% of variable run costs, while human oversight accounts for 70โ€“75% (McKinsey QuantumBlack, 2026). The stage your business case treats as free overhead is the majority of the run cost.

Second, it degrades under load. A survey of 2,500 knowledge workers found 42% spend more time verifying AI output than they save using it, and 52% regularly correct AI-generated work produced by colleagues (Adaptavist, 2026). Push more volume through an unchanged review stage and the reviewers absorb it as rework โ€” which is exactly how a 240% upstream gain converts into a 30% downstream one.

So the weak link is not merely slow. It is the cost centre, and it is already saturated.

The Honest Counter

Three limits, stated before someone else states them.

The authors have a disclosed relationship with a vendor. Demirer and Musolff both previously held postdoctoral research positions at Microsoft and now work as paid research consultants for the company โ€” disclosed on the working paper itself. The finding cuts against the commercial interest (it caps the claimable output effect of coding tools), which is the direction that makes a conflict less worrying. Note it anyway.

The estimates moved between drafts, and that matters for how you cite them. The May 2026 version reported a smaller sample and different coefficients; the September 2026 revision reports more than 500,000 developers, the 30/180/240% generational effects, and the 0.23 elasticity. The attenuation pattern held across both. If you quote a figure from this paper in a board deck, quote the current revision and date it โ€” working papers are not settled results, and the ones worth citing are the ones whose shape survives revision.

One industry is not every industry. Software has unusually clean stage boundaries and unusually good telemetry. Your finance close or your fulfilment chain may have deeper human involvement at every stage, which would compress both the upstream gain and the attenuation. The direction transfers. The magnitudes are not yours until you measure them.

What to Decide This Quarter

Four moves. Three cost nothing but attention.

  1. Name the last human gate in one process. Not the process owner โ€” the specific approval, review or sign-off step that every unit of work must pass through before it counts as delivered. If you cannot name it in a sentence, that is the finding.
  2. Measure the two stages separately. Your current metric almost certainly counts activity at the accelerated stage: drafts produced, tickets triaged, quotes generated. Add one counter at the delivery stage. The ratio between them is your attenuation, and it is the only number in this article that is actually about your company.
  3. Move the next increment of AI budget downstream. If the weak-link result holds in your pipeline, the marginal return on another upstream seat is near zero and the marginal return on unblocking the gate is whatever that gate is currently holding back. That is a reallocation, not new spend.
  4. Redesign the gate before you widen the funnel. Raising upstream volume against an unchanged review stage produces queue and rework, not output. Sequence it the other way: fix the gate, then let the volume through.

Your AI productivity gains are real. The study says so, at a scale no internal pilot will ever reach.

They are also sitting in a queue behind a human being nobody budgeted for. The question for this quarter is not how much faster your team can produce work. It is how much of that work your organization can still finish.

Ready to go beyond the CV?

Scovai's AI-powered Talent Passport reveals what resumes can't: personality, potential, and true job fit.