Scovai Scovai
AI & Operations 2026-09-14 1 min read

You're Optimizing the Cheap Quarter of Your Agent Bill: McKinsey Puts Human Oversight at 70โ€“75% of an AI Agent's Variable Cost

DSL

Dr. Sarah Liu

You're Optimizing the Cheap Quarter of Your Agent Bill: McKinsey Puts Human Oversight at 70โ€“75% of an AI Agent's Variable Cost

Price a customer service agent in banking and the token bill comes to 20 to 25 percent of its variable run cost. Human oversight โ€” functional and risk experts checking the agent's work โ€” comes to 70 to 75 percent (McKinsey QuantumBlack, 2026).

Same workflow. Same invoice. And almost every agent business case I have read this year optimizes the smaller number.

Model selection, prompt caching, a cheaper inference tier for low-stakes calls โ€” that is where the engineering attention goes, because that is where the cost is legible. It arrives monthly, itemized, attributable. Oversight arrives as salaried hours inside someone else's cost centre, so it never gets counted as AI agent cost at all. It gets counted as the reason the pilot quietly failed to scale.

What McKinsey Actually Priced

QuantumBlack's August 24, 2026 guide to the economics of agentic workflows does something most agent analyses skip: it prices the completed work, not the software. Agents, deterministic systems, and the humans reviewing the output, all in one fully loaded number.

The headline case is bank account opening. The workflow runs five to seven agents alongside multiple deterministic systems, with two to four teams of people providing oversight. Fully loaded cost per completed onboarding falls from roughly $50โ€“$150 to about $10โ€“$30 (McKinsey QuantumBlack, 2026).

That is a three-to-five-fold improvement. This is not an argument that agents don't pay. It is an argument that you are tracking the wrong line item to find out whether yours do.

For scale: McKinsey puts customer-facing single-agent workflows at some banks at $20,000โ€“$30,000 a year to run, and a multiagent team at $100,000โ€“$200,000. Those figures come from McKinsey's analysis of public research and public pricing rather than audited client ledgers โ€” so hold the ratios firmly and the absolute dollars loosely.

What Actually Drives AI Agent Cost: The Exception Rate, Not the Model

In banking customer onboarding, McKinsey expects 10 to 20 percent of agentic runs to be reviewed by risk and functional experts.

That review rate is the entire ballgame, and it is worth walking through the arithmetic slowly โ€” the illustration below is mine, not McKinsey's, but the inputs are theirs.

Take a workflow at a 15% review rate, with 20 minutes of expert time per review. Across 1,000 runs that is 50 hours of senior functional time. Drive the review rate down to 5% and the same 1,000 runs cost about 17 hours. You have removed two-thirds of the expensive line without touching a model, a prompt, or a vendor contract.

And note who is doing that reviewing. McKinsey specifies functional and risk experts โ€” the people whose judgment the workflow exists to protect, not a junior QA tier you can staff cheaply. Oversight cost is expensive precisely because cheap oversight is not oversight; the reviewer has to be senior enough to overrule the agent. Any plan that assumes the review layer can be delegated downward is really a plan to stop reviewing.

Now compare that to the lever everyone actually pulls. McKinsey identifies only three drivers of token cost: model quality, latency requirement, and task frequency. Frontier pricing in early July 2026 ran about $5 per million input tokens and $30 per million output for a top-tier model, against $0.20 and $1.25 for an earlier lightweight one. A twenty-five-fold spread โ€” applied to a quarter of the bill.

And here is the trap inside the obvious optimization. Downgrade the model to cut tokens and you raise the error rate. A higher error rate raises the exception rate. A higher exception rate raises the reviewer hours, which is the three-quarters. Token savings routinely finance an oversight increase that nobody measures, and the workflow gets more expensive while the dashboard shows it getting cheaper.

Why the Reviewed Hours Don't Disappear

There is independent evidence that the oversight line is structural rather than a maturity problem you outgrow.

MIT Technology Review Insights and Microsoft surveyed 300 technology executives and practitioners and ranked 101 agentic tasks on a 0โ€“100 trust scale. Confidence tracked task verifiability and business-context completeness, not model capability: automated report generation scored 83.5 and boilerplate code 82.5, both with a single objective grading metric, while service-mesh configuration scored 37.5 and disaster-recovery testing 43 โ€” the same underlying models, minus a clean way to tell whether the output was right. Accountability worried 48% of respondents and hallucinations 47%, and 59% already plan permanent human oversight (MIT Technology Review Insights, 2026).

Production data points the same direction. A study of Perplexity's agent product against its search product found matched-task completion falling from 269 minutes to 36 โ€” an 87% time and 94% cost reduction โ€” while the follow-up work users did shifted upward, into verification and extension (arXiv 2606.07489, 2026).

Human work relocates to the review step. McKinsey is simply the first to put a price on where it lands.

Volume and Reuse Decide Whether the Math Works

The other half of McKinsey's economics is fixed cost, and it is the half that determines whether mid-market agent programs clear the bar at all.

A conversational agent onboarding 2,500 new customers a year costs $10,000โ€“$15,000. Double the volume and the cost rises only to $15,000โ€“$20,000 โ€” because AI infrastructure and agent orchestration are largely fixed, and fixed costs amortize. Note what sits inside that orchestration line: standing data-science capacity to maintain the agent in production, which McKinsey observes "often need[s] to be tweaked every couple of days."

Low-volume workflows therefore lose twice. They get no amortization on the fixed cost, and their oversight ratio stays high because they never accumulate enough runs for anyone to find the exception patterns and design them out.

The mid-market escape route is reuse rather than scale. If no single workflow carries enough volume to amortize the infrastructure and orchestration line, the question becomes how many workflows can run on one agent build โ€” McKinsey's framing is build once, deploy across workflows. Three adjacent processes sharing an agent, a context layer, and a review protocol clear the fixed-cost bar that none of them clears alone. That is an architecture decision made before the first pilot, and it is very difficult to retrofit once three teams have each bought their own point solution.

That asymmetry is visible in the adoption data. McKinsey's State of AI 2026 found organizations above $1 billion in revenue scaling agents rising from 27% to 40%, while organizations under $1 billion stayed flat at 22%. Around 20% reported that AI operating costs constrained their use, and the share attributing at least some EBIT impact to AI held at 37%, unchanged year over year (McKinsey, 2026).

Adoption compounds. Reported financial impact does not. That is the signature of a cost line nobody is managing.

The Honest Counter: This Is Banking, and It's an Estimate

Two caveats you should hear from me rather than from your CFO.

First, the 70โ€“75% figure is drawn from regulated, customer-facing financial services. Oversight intensity is a function of regulatory exposure and error consequence. Internal ticket triage at a 200-person logistics firm sits on a different curve, and the split there may be closer to even.

Second, these are modelled costs built on public pricing and public research, not measured ledgers from a client engagement. Treat them as a well-specified estimate, not a measurement.

But notice which way the error runs. Token prices have fallen steadily and will keep falling. Reviewer salaries have not and will not. Every month that passes moves the ratio further against tokens, not back toward them. If 70โ€“75% is wrong for your context today, the honest correction is to ask what it becomes in eighteen months โ€” not to assume it collapses back to something the token dashboard can see.

Price the Reviewer Hours Before You Approve the Pilot

Five things to change in how you evaluate the next agent proposal.

  1. Put the exception rate in the business case as a named number. Percentage of runs needing human review, expected minutes per review, loaded hourly cost of the reviewer. If no one will commit to those three figures, you do not have a business case โ€” you have a demo with a budget attached.
  2. Engineer verifiability before you widen scope. The MITโ€“Microsoft ranking shows trust follows tasks with a single readable success metric. That property is buildable: instrument the metric, encode the business context, narrow the scope. Sequence deployments by verifiability, not by which task looks simplest.
  3. Rank candidate workflows by annual run count first. Reuse and volume are what amortize the fixed cost. A clever low-volume use case is a worse investment than a dull high-volume one, every time.
  4. Budget the maintenance line explicitly. Agents tweaked every couple of days require standing engineering capacity. That is a permanent fixed cost, not a project expense.
  5. Report completed-work ROI. Fully loaded cost to finish the job โ€” humans, agents, and deterministic systems together โ€” measured against the value of the job. Not cost per token, and not hours notionally saved.

McKinsey's own prescription is the inverse of the instinct: rather than optimizing model selection, redesign the workflow to cut exception rates and simplify the review processes that trigger costly human intervention. That is a process engineering mandate sitting in what everyone has been treating as a procurement decision.

What to Decide This Quarter

Find the agent pilot currently waiting on your signature and ask for one number that is not in the deck: expected reviewer hours per hundred runs.

If the answer is a shrug, you are not being asked to approve a piece of software. You are being asked to staff a review team that nobody has budgeted โ€” and to do it at the three-quarters of AI agent cost that your dashboard was never built to show you.

Ready to go beyond the CV?

Scovai's AI-powered Talent Passport reveals what resumes can't: personality, potential, and true job fit.