Scovai Scovai
AI & Operations 2026-09-17 1 min read

Chat Logs Aren't a Workforce Study: A New Fed-Harvard Paper Says the AI-Usage Benchmark Under Your Business Case Counts Work Your Jobs Don't Contain

DSL

Dr. Sarah Liu

Chat Logs Aren't a Workforce Study: A New Fed-Harvard Paper Says the AI-Usage Benchmark Under Your Business Case Counts Work Your Jobs Don't Contain

More than 15% of OpenAI chats are classified as "Edit written materials or documents." Only 2.4% of workers hold an occupation that contains that task in O*NET (Bick, Blandin, Deming & Schumacher, NBER, 2026). That gap sits inside the AI usage benchmark under most 2027 operating plans.

Now the other direction. The single task that captures the largest share of genAI use in a nationally representative survey is "Direct organizational operations, activities, or procedures" โ€” 4.1% of all reported uses, held by 43% of workers. Its share in the OpenAI data is 0.0%. In Microsoft's, 0.0%. In Anthropic's, 0.2%.

The work most characteristic of operations is the work the chat logs cannot see at all. If your 2027 AI plan was sized against one of those vendor usage studies โ€” and most mid-market plans were, because those studies are free, current, and enormous โ€” you are reading a map with your own territory left off it.

What This Survey Did That the Chat Studies Can't

Alexander Bick (Federal Reserve Bank of St. Louis), Adam Blandin (Vanderbilt), David Deming (Harvard Kennedy School) and Tyler Schumacher (Vanderbilt) built the first task-level genAI adoption index from a nationally representative survey rather than from platform telemetry (NBER Working Paper 35677, 2026).

The mechanism matters. They pooled four waves of the Real-Time Population Survey โ€” 13,920 employed respondents aged 18 to 64 โ€” and linked each person's reported genAI use to their detailed occupation and to the specific ONET work activities that occupation actually contains. Because they know the respondent's job, they can attribute a use to a task within a known job architecture*.

A chat classifier cannot do that. It reads the text of a conversation and typically has no idea what the user does for a living. So it does the only thing available: it assigns the chat to a task title that describes the visible action. Editing. Gathering information. Designing a system.

That is not a bug in anyone's classifier. It is a structural limit of the input, and the paper names it plainly: chat classifiers "may be predisposed to assigning chats to task titles describing basic, generic actions."

How Far the AI Usage Benchmark Drifts

The consequence is extreme concentration. Ten tasks out of 332 account for 53.0% of use in the OpenAI data, 46.0% in Anthropic's, and 60.8% in Microsoft's โ€” against 22.2% in the survey (Bick et al., NBER, 2026). The largest single task takes 15.1%, 16.7% and 23.2% of the three chat datasets respectively, versus 4.1% in the survey.

Then the number that should end the debate about whether these are two views of one reality. The correlation between the survey's task shares and the chat datasets' task shares is 0.11 for OpenAI, 0.34 for Anthropic, and 0.10 for Microsoft. Two of the three are statistically almost unrelated to how workers say they use the tools across their actual tasks.

The Second-Order Error Is the One in Your Plan

Direct usage share is not usually what lands in a business case. The chain is longer, and the paper follows it: chat-task classifications get aggregated into genAI occupation shares, and those occupation shares then circulate as proxies for AI exposure by role โ€” the input behind rollout sequencing, training budgets and headcount forecasts. The authors name four published measures built this way and conclude the proxies "are likely inaccurate and offer a distorted view."

So when a deck tells you which of your functions are most AI-exposed, ask where the ranking came from. If it traces back to chat logs, it inherits a classifier that never knew what anyone's job was โ€” and its characteristic error is to over-rank work that looks generic in a transcript.

Your Other Input Is Also Weaker Than It Looks

The obvious correction is to drop usage benchmarks and use a proper occupational exposure score instead. The paper tested that too, against the two most-cited indexes in the literature.

Regressed on actual adoption, the Eloundou et al. (2024) measure yields an adjusted Rยฒ of 0.081 and the Felten et al. (2021) measure 0.086 โ€” roughly half of the occupation-level variation available to be explained (Bick et al., NBER, 2026). The scores are genuinely predictive. They are not close to sufficient.

And there is a harder constraint underneath. Occupation itself explains less than 20% of the variation in worker-level adoption (Bick et al., NBER, 2026). Whatever is driving who actually adopts sits mostly inside the job title, not between job titles. Two analysts on the same team, same tools, same tasks โ€” one has rebuilt their week around the tool and the other opens it twice a month.

That is a planning problem with a specific shape. Any allocation you make per role โ€” seats, training hours, enablement sequencing โ€” is being made on a variable that accounts for under a fifth of the outcome you care about.

It also explains a pattern most operations leaders have already seen and filed under something else. A function comes back from the same enablement session with wildly different results; the reflex reading is motivation, or manager quality, or resistance to change. The adoption data says a large share of that spread was always going to be there, because it is person-level and the intervention was role-level. Running the session again, harder, does not address it.

"Broad But Shallow" Is a Precise Claim, Not a Hedge

The adoption picture the survey produces is easy to misquote in either direction. The precise version:

  • As of May 2026, 45% of workers use genAI for work; the authors treat this as a lower bound, since embedded and passive uses often go unreported.
  • More than 80% of occupations and 40% of tasks show adoption above 20%.
  • But only 2.8% of tasks exceed 50% adoption, and none exceed 70%.

Read those together. There is no task in the American economy where genAI use is near-universal. The occupations that do show very high adoption โ€” roughly 15% clear 70% โ€” get there not through one saturated task but through a collection of moderately-adopted ones.

Independent data agrees on the shape. Google's ATLAS study, built from 14.65 million de-identified interactions rather than self-report, found the median occupation using Gemini on about 21% of its tasks, with end-to-end automation intent under 10% of non-routine cognitive conversations (Google ATLAS v1.0, 2026). Behavioral telemetry across 120,620 workers put just 2% at workflow-integration maturity โ€” using AI inside a restructured process rather than as a side lookup (ActivTrak Productivity Lab, 2026).

Three methods, three datasets, one conclusion: wide coverage, thin penetration. Which means a plan that models an AI-exposed role as an AI-transformed role is off by most of the distance.

The Honest Counter

Three limits worth holding.

The 2.4% figure is partly an ONET artifact, and the authors say so. Their exact words: "the share of workers who perform editing is clearly much larger than 2.4%." ONET embeds editing inside higher-order tasks like preparing reports or drafting legal documents. So the gap between 15% and 2.4% is not pure classifier error โ€” some of it is the reference taxonomy declining to name an activity that is genuinely everywhere. The direction of the distortion is well established; the magnitude of any single pair of numbers is softer than it looks.

Self-report has its own failure modes. The survey knows what people say they do with the tools. Chat logs know what actually passed through the window, at a scale no survey will ever reach. Neither is the ground truth; they are answering different questions, and the paper's contribution is to show how far apart the answers are, not to declare one of them fake.

A national survey is not your company. The RPS describes the US workforce. Your task inventory, tool stack and job design are not the national average, and nothing here tells you which of your roles will adopt.

That last limit is the one that actually matters โ€” and it points in a single direction. Every external benchmark in this space, chat-log or survey, is a prior. None of them is a measurement of your organization.

What to Decide This Quarter

Four moves. None requires new spend.

  1. Trace every AI usage benchmark in your plan back to its source. If a figure originates in platform chat logs, mark it as a directional prior, not an input to a headcount or budget model. This is a thirty-minute audit of a document you have already written.
  2. Build one task inventory for one function. Twenty to thirty real activities for a single team, written in your language rather than O*NET's. It is the only artifact that lets you ask "does the tool touch this?" instead of "is this role exposed?" โ€” and it survives every vendor benchmark revision, because it describes your work rather than someone else's traffic. Budget an afternoon with the team lead, not a consulting engagement.
  3. Stop allocating enablement by role. With occupation explaining under 20% of worker-level adoption, role-based rollout is close to a coin flip. Allocate against observed use instead, and let the first cohort be whoever is already ahead in your own telemetry.
  4. Set your target at the task, not the seat. Pick one task in that inventory and drive it past 50% adoption. Fewer than three in a hundred tasks nationally have crossed that line โ€” which is exactly why crossing it deliberately is a defensible thing to report.

Your AI usage benchmark answered a question you never asked: what do chats look like. The question on your desk is what your people do all day, and whether the tool has reached any of it.

One of those questions has a published answer. The other one only has yours.

Ready to go beyond the CV?

Scovai's AI-powered Talent Passport reveals what resumes can't: personality, potential, and true job fit.