Scovai Scovai
AI & Operations 2026-08-27 1 min read

AI's Savings and Its Rework Land on Different Desks: 52% of Knowledge Workers Now Correct Their Colleagues' AI Output

DSL

Dr. Sarah Liu

AI's Savings and Its Rework Land on Different Desks: 52% of Knowledge Workers Now Correct Their Colleagues' AI Output

Fifty-two percent of knowledge workers now regularly correct AI-generated work produced by someone else (Adaptavist, 2026). Not their own output โ€” a colleague's. In the same survey of 2,500 workers across five countries, 42% report spending more time verifying AI output than they save by using it. Read those two figures together and you have the reason your AI rework never shows up in any business case: the person who banks the saving and the person who absorbs the cost are not the same person, and frequently not in the same team.

That is an accounting problem before it is a technology problem.

The Study, and What Makes It Different

Adaptavist fielded the research in March 2026 across the UK, US, Canada, Germany and Spain โ€” 2,500 knowledge workers, published as a follow-up to its 2025 work on digital transformation (Adaptavist, 2026). The report coins a useful term: the verification tax.

The headline numbers:

  • 52% regularly correct AI-generated work from colleagues
  • 42% spend more time verifying AI output than they save using it
  • 49% say poor-quality AI output actively slows projects down
  • 46% say AI has made their work feel more repetitive and less meaningful

The last figure is the one most people skip past. Correcting someone else's machine output is not interesting work. It is not developmental, it rarely gets credited in a review, and it is invisible to whoever approved the tool.

Note also what the study does not show. It is not an anti-AI dataset: 67% of the same respondents want their organisation to increase AI use, 60% feel adequately trained, and 66% say adoption was communicated transparently (Adaptavist, 2026). These are people who want the technology and are describing a distribution failure, not a capability failure.

Why Your Dashboard Says Yes While Your Cycle Times Say Nothing

Most mid-market AI measurement is per-seat. Licences deployed, prompts run, self-reported hours saved, adoption percentage by department. Every one of those metrics is collected at the point of generation.

The rework is collected nowhere, because it happens at the point of receipt.

Marketing generates a first draft in ninety seconds and books ninety minutes saved. Legal spends forty minutes fixing a citation that was confidently wrong. Marketing's dashboard is honest. Legal's forty minutes lands in a general workload line, and nobody reconciles the two. Repeat this across finance, ops, support and product, and you get the pattern that has become almost universal in mid-market reporting: reported time savings rise every quarter, and cycle times stay exactly where they were.

Three properties that make this cost hard to see

It crosses a boundary. Cost accounting is organised by team. A cost that originates in one team and lands in another has no natural owner, so it is not tracked by either.

It is unbilled seniority. The corrector is usually the more senior party โ€” the person with enough domain judgment to spot what is wrong. You are paying your highest hourly rates to clean up output generated at your lowest.

It is socially awkward to report. "I spent my afternoon fixing a colleague's AI draft" is not a comfortable line in a status update, so it gets logged as the underlying task instead.

This Is Not the Clawback You Have Already Read About

Two adjacent findings get conflated with this one, and the distinction determines which lever you pull.

Glean's Work AI Index โ€” 6,000 digital workers, produced with researchers from Stanford, UC Berkeley and Harvard โ€” found AI saves roughly 11 hours a week while workers burn about 6.4 hours "botsitting": supervising, debugging and tool-switching (Glean, 2026). That is a same-worker clawback. The person who saves the time pays it back themselves, and the fix is context access โ€” what your systems let the model reach.

METR's randomised trial pointed at a different failure: 16 experienced open-source developers, 246 real tasks in their own repositories. They forecast a 24% speed-up, believed afterwards they had gained 20%, and were measured 19% slower (METR, 2025). That is a perception failure โ€” self-reported savings are unreliable even from experts working on their own code.

Adaptavist describes a third thing: a lateral transfer. Not the same worker paying themselves back, not a misperceived gain, but a cost that leaves one cost centre and arrives in another. Context tooling will not fix it. Better self-reporting will not fix it. Only measuring at the boundary will.

The three findings compound, incidentally. If savings are overstated at source (METR), partially reabsorbed by the same worker (Glean), and partially exported to a colleague (Adaptavist), the residual net gain in a typical mid-market workflow is a much smaller number than any of the individual studies imply.

Sizing It: What the Verification Tax Looks Like at 200 FTE

The study reports proportions, not hours, so any euro figure is your arithmetic rather than Adaptavist's. Do it anyway โ€” the order of magnitude is the argument.

Take a 200-person company where 120 people do knowledge work. If 52% regularly correct colleagues' AI output, that is roughly 62 people. Assume โ€” conservatively, and this is the assumption to test rather than trust โ€” two hours a week each. That is 124 hours weekly, about 6,400 hours a year, or the equivalent of three full-time roles you never hired, budgeted, or discussed.

Now note who those hours belong to. Correction requires the judgment to know what is wrong, so it concentrates in your senior individual contributors โ€” the same people you would name if asked who is hardest to replace. You are spending three FTEs of your scarcest judgment on cleanup that is not on anyone's objectives.

The number is illustrative. The exercise is not: run it with your own headcount and your own estimate, put it in front of whoever owns the AI budget, and watch how quickly the conversation shifts from licence count to workflow design.

What the study cannot tell you

Three limits, stated plainly. It is self-reported, so the 42% is a perception of net time, not a measurement of it โ€” and METR's work is a direct warning about how far self-reports drift. It is cross-sectional, so it captures March 2026 and cannot say whether the tax is growing or decaying. And it surveys workers across five countries and all company sizes, not mid-market operations specifically; the direction should transfer, the magnitude is unestablished in your context.

What survives all three is the structural point: correction crosses a boundary that measurement does not.

The Obvious Objection, and Where It Holds

The reasonable counter: this is early-adoption friction. Output quality improves, review norms mature, and the tax decays. Some of that is certainly true. First-generation usage in any tool class carries a correction premium that falls with familiarity.

Two things constrain the optimism.

First, the correction is not primarily about model quality โ€” it is about accountability placement. As long as generating is cheap and reviewing is expensive, volume rises to the level the generator finds convenient and the reviewer absorbs the difference. Better models raise the volume as fast as they raise the quality.

Second, there is evidence the friction is being formalised rather than dissolved. BambooHR's 2026 workforce research found 81% of leaders claiming AI productivity gains while 49% of workers said AI had delivered no value at all (BambooHR, 2026). When unverified gains are already priced into leadership expectations, the correction work does not get resourced โ€” it gets absorbed silently by the people whose performance is being measured against the claimed gain.

So: the friction argument holds for the magnitude of the tax. It does not hold for its direction.

Measuring AI Rework at the Handoff Boundary

Four moves. All are measurement or policy changes; none require new spend.

1. Instrument the handoff, not the keystroke. Pick your two highest-volume cross-team workflows and record three things on the receiving side: review hours, rejected drafts, and rounds-to-accept. You already have the review step. You are simply not timing it. A month of this produces a real net-hours figure instead of a self-reported one.

2. Assign review ownership before a tool goes live. Every AI-assisted workflow that crosses a team boundary needs a named reviewer and a budgeted number of hours. If no one will fund the review, the workflow is not ready โ€” you are proposing to route an unbudgeted cost into a colleague's week.

3. Make provenance a norm, not a confession. Requiring a one-line note on AI-assisted work that crosses a boundary โ€” what was generated, what was checked, what wasn't โ€” costs the sender seconds and saves the receiver the reconstruction. It also, quietly, restrains volume, because unreviewed bulk output becomes visibly attributable.

4. Rebalance the sample of who you ask. Adoption surveys almost always poll the generators. Poll the receivers: how much of your week is spent correcting AI-assisted work you did not produce? That single question, asked of your senior individual contributors, will tell you more about your true AI ROI than your entire licence-utilisation dashboard.

Note the pattern in all four: none of them slows adoption. They relocate the measurement to where the cost actually lands.

One Decision for This Quarter

Your AI rework is not a rounding error, and it is not invisible because it is small. It is invisible because your instrumentation stops at the moment of generation, and the cost begins one desk over.

So the question is not whether your AI is delivering value. It is narrower and far more answerable: in your two highest-volume cross-team workflows, who is currently doing the correcting โ€” and has anyone ever asked them how long it takes?

If you cannot name that person by the end of the week, you do not have an AI ROI problem. You have a measurement problem wearing an ROI costume, and it will keep telling you exactly what you want to hear until you move the meter to the other side of the handoff.

Ready to go beyond the CV?

Scovai's AI-powered Talent Passport reveals what resumes can't: personality, potential, and true job fit.