Three economists at the St. Louis Fed ran roughly 490,000 earnings-call transcripts from 5,198 US public firms through a language model, tagged 910,955 sentences about productivity, and sorted them by tense. About 95% of the AI-related productivity statements referred to gains that had not happened yet. For non-AI productivity talk, the figure was around 75% (St. Louis Fed, 2026).
Over the same window, utilization-adjusted total factor productivity across the US economy grew 0.07% in the four quarters ending Q1 2026.
That is the whole problem in two numbers. The public record of AI productivity is almost entirely a record of expectations, and your AI business case is being underwritten against it as if it were a record of results.
The Benchmark You Are Using Is a Tense, Not a Number
Start with what the study actually did, because the method is what makes the finding hard to wave away.
Serdar Ozkan, Aakash Kalyani and Nicholas Sullivan pulled S&P Global transcripts covering 2000 to 2025, keyword-matched candidate sentences, then used a Qwen3-3B classifier to label each one on three axes: is it about AI, is productivity described as rising, falling or unchanged, and is it framed in the past, present or future. Non-productivity sentences were dropped. What survived was 910,955 tagged statements (St. Louis Fed, 2026).
Two distributions came out of it, and both are lopsided in the same direction.
On tense: ~95% of AI productivity sentences are forward-looking, against ~75% for productivity talk generally. On direction: ~95% of AI-related sentences describe productivity as rising, against ~75% for non-AI. The share describing an AI-driven productivity decrease is negligible.
The obvious objection is early-hype lag โ of course 2023 was aspirational, the tools were new. The data does not support that read. The forward-tense share is stable across both 2023 and 2025, even as AI's share of all productivity discussion climbed from roughly zero pre-ChatGPT to about 15% by the end of 2025 (Fortune, 2026). Two full years of deployment did not move the conversation from will to did.
So when a peer company's AI productivity claim reaches you โ in a conference keynote, a vendor case study, a competitor's investor deck โ the base rate says it is a forecast. Not a lie, not necessarily wrong. A forecast, wearing a benchmark's clothes.
The Same Gap Shows Up Inside the Firm
If the earnings-call finding were only about investor-facing language, you could discount it as capital-markets theatre. It isn't. The Atlanta and Richmond Fed CFO Survey found the identical wedge in private, non-promotional responses.
Surveying more than 700 corporate executives in March 2026, the researchers asked firms to self-report AI's effect on output per worker: the answer was 1.8% for 2025, with 3.0% expected for 2026. The team then computed what the same firms' own numbers implied โ AI-attributed revenue against AI-driven employment change โ and got 0.4% to 0.8% in manufacturing, construction and low-skill services, rising to about 2.2% in finance (Atlanta Fed, 2026).
Executives overstated their own realized gains by a factor of two to four, using their own data, in a confidential survey, with no audience to impress.
That reframes the earnings-call finding. This is not primarily a disclosure problem or an IR-discipline problem. It is a measurement problem, and it exists at the level of the individual operating executive. The authors reach for Solow's 1987 line about seeing the computer age everywhere except in the productivity statistics, and the parallel is earned: two independent Fed measurements โ one textual across 5,198 firms, one survey-based across 700-plus executives โ converge on the same conclusion from opposite methodological directions.
Why Forecast Contamination Is Worse in the Mid-Market
A 50โ500 FTE operations function is exposed to this asymmetry more than a large enterprise, for a structural reason that has nothing to do with sophistication.
The CFO Survey's size split is stark. Around 60% of firms invested in AI in 2025 โ roughly 80% of large firms, about half of small ones โ and over 80% expect to invest in 2026. But the amounts diverge by orders of magnitude: about 60% of smaller firms plan to spend under $20,000 on AI in 2026, against 14% of larger firms, while roughly 30% of large firms plan to spend over $1 million, against 1% of smaller firms (Richmond Fed, 2026).
A firm spending $700,000 can afford to instrument the deployment: baseline the process, run a holdout, measure cycle time before and after. A firm spending $18,000 cannot, and does not. It buys the licences and infers the result from how the team feels about the tool.
Which means the mid-market is disproportionately reliant on external benchmarks precisely because it cannot generate internal ones. And the external benchmark pool is 95% forecast. The less measurement capacity you have, the more weight you put on the least reliable evidence class available. That is the trap, and it tightens as budgets get smaller.
There is a second-order effect worth naming. When peer claims are optimistic and your own measurement is thin, the gap between your observed result and the benchmark reads as your execution failure. Ops leaders then respond by buying more tooling or restructuring the rollout, when the honest diagnosis is that the comparison was never valid. You are not behind. You are measured, and they are projected.
The headcount forecast is the same shape
The workforce numbers in the CFO Survey deserve separate attention, because they are the ones most likely to be sitting in your plan as if they were settled.
Aggregate employment effects from AI were close to zero. Large firms expect a 0.8% headcount reduction in 2026; the smallest firms expect roughly a 1% increase. On composition, executives project the routine clerical share of employment falling 0.76% by 2026 and 2.19% by 2028, partly offset by growth in skilled technical roles โ and smaller companies are more likely to expand technical headcount than to cut anywhere (Richmond Fed, 2026).
Note the tense again. Every one of those figures is an expectation about 2026 and 2028, produced by the same executives who overstated their realized 2025 productivity gains by two to four times. The bias does not politely confine itself to the productivity question. If your 2027 operating plan assumes a clerical reduction because that is what the survey data shows, you have imported a forecast from a population with a documented, measured optimism problem โ and booked it as a planning assumption.
What the Evidence Does Not Say
Three limits, because a business case built on a misread of this study fails the same way as one built on the hype.
Forward-looking is not the same as false. A sentence in the future tense is a statement about intention and expectation, not a fabrication. Some of those forecasts will land. The finding is about the evidentiary class of the claim, not its truth value โ and the correct response is to reweight, not to dismiss.
Aggregate TFP is a blunt instrument. The 0.07% figure is economy-wide and slow-moving; measured productivity has historically lagged general-purpose technologies by years, and firm-level gains can be real while the national statistic sits flat. The TFP number is context for the forecast asymmetry, not proof that no one is getting anything.
Classification is imperfect. Tense-tagging by a small language model over nearly a million sentences carries error, and earnings-call language is hedged by construction โ legal review pushes executives toward forward-looking framing on any uncertain topic. Some of the 95% is a disclosure artefact rather than an absence of results.
None of that rescues the benchmarking practice. If anything, the CFO Survey tightens the case: strip out the lawyers, the analysts and the incentive to impress, and executives still overstated realized gains by two to four times against their own figures. The asymmetry survives every deflation you can reasonably apply to it.
The Three Fields That Fix Your AI Business Case
This is a documentation change, not a strategy change, and it costs an afternoon.
Take every AI initiative currently in your budget or pipeline. For each one, require three fields before it advances:
One realized metric. Not a projected saving, not a per-seat time estimate from the vendor. A specific operational number: average handling time, first-pass yield, days to close, tickets resolved per FTE. If nobody can name one, the initiative is a hypothesis and should be funded and reviewed as an experiment, not as a return.
Its pre-AI baseline. The value of that metric before the tool arrived, with the measurement window stated. Missing baselines are the single most common failure โ teams start measuring the day the pilot goes live, which makes any subsequent number uninterpretable. If the baseline is gone, reconstruct it from historical data before you spend another quarter.
The date it was measured. Realized results have a measurement date attached. Forecasts do not. This field alone sorts your portfolio faster than any framework, because a claim that cannot produce a date is not a result.
Then apply the same test to every external number you are benchmarking against. Vendor case study, peer keynote, competitor deck: does it name a realized metric, a baseline, and a date? The St. Louis Fed's finding predicts that about nineteen in twenty will fail on at least one. Those go in a folder marked expectations, and they stop influencing budget.
The discipline is unglamorous, and it will make your AI programme look worse on paper this quarter. That is the point. Right now you are comparing your measured results to everyone else's projections and calling the difference underperformance. Fix the comparison and you will find you are not behind โ you are simply one of the few people in the room actually counting.
Before your next AI investment decision, ask one question of every number in the room, internal or external: when was this measured? Everything that cannot answer is a forecast, and forecasts do not belong in a business case's evidence column.