A June 2026 systematic review and meta-analysis of 19 studies and 61 effect sizes puts a number on something most operations leaders have sensed but never measured: when people create with generative AI, their outputs converge. The pooled homogenization effect is d = 0.334 โ small, robust, and not explained by publication bias (de Rooij & Biskjaer, 2026).
"Small" is the word that gets this finding filed and forgotten. It shouldn't be.
Individually, the work usually gets better. Collectively, it gets more alike. Those two facts are not in tension โ they are the same mechanism viewed from two altitudes. Only one of them appears on your productivity dashboard.
For an operations function that standardized on a single-vendor AI rollout last year โ which describes most companies between 50 and 500 people โ the practical question is not whether to keep using it. It is whether anyone has checked what happened to the spread.
What d = 0.334 Actually Means at Team Scale
Cohen's d of 0.33 is a modest shift between two distributions. On a single document, it is invisible. You would not catch it comparing two proposals side by side, and no reviewer in your company will ever flag it.
The meta-analytic framing is what makes it operationally relevant. The authors describe the finding not as a drop in quality but as a reorganization of creative diversity โ subtle in isolation, consequential at scale (de Rooij & Biskjaer, 2026). This is a variance finding, not a quality finding. That distinction matters, because every measurement system in a mid-market operations function is built to track means and almost none of them track spread.
A March 2026 opinion review in Trends in Cognitive Sciences traces the same pattern to its source. Sourati and Dehghani argue that as billions of people route their writing and thinking through the same handful of large language models, those models standardize how people express and structure ideas โ LLM outputs are measurably less varied than human writing and skew toward WEIRD (Western, educated, industrialized, rich, democratic) norms (Sourati & Dehghani, 202600003-3)). Their concern is the erosion of cognitive diversity, which is the raw material of both creativity and problem-solving.
Restate that in operating terms. Cognitive diversity is not a values statement. It is the input variance that makes a portfolio of options worth generating in the first place.
The Mechanism Is Anchoring, Not Laziness
The instinct in most leadership teams is to read convergence as effort collapse โ people are taking the first draft and shipping it. The experimental evidence points somewhere less flattering to our controls.
In a controlled online experiment on short-story production, writers who received story ideas from an LLM produced work rated as more creative, better written, and more enjoyable โ with the largest gains among the less creative writers. And yet the AI-enabled stories were more similar to each other than the stories written by humans alone (Doshi & Hauser, 2024). Both effects came out of the same intervention.
The authors name the structure explicitly: it resembles a social dilemma. With generative AI, writers are individually better off, but collectively a narrower scope of novel content gets produced.
A separate study of human ideation reaches the compatible conclusion from the input side. Sessions where participants brainstormed with an LLM produced idea sets that were significantly less semantically diverse than sessions without one โ the individual ideas were fine, the set was narrower (Anderson, Shah & Kreminski, 2024).
That is anchoring, and it operates before anyone has decided how hard to work. The model supplies a starting point; the human elaborates from it. Elaboration preserves quality and destroys independence.
Why no one in your company will stop
Read the social dilemma as an incentive map, because that is what it is.
The gain from the AI-assisted draft is captured by the individual โ faster turnaround, cleaner prose, a better performance conversation. The loss is charged to the collective โ a proposal set that reads like every competitor's, a hiring rubric indistinguishable from the sector template. No single contributor experiences the loss, and no single contributor has any reason to stop.
Which means this is not a behavior problem. It is an org-design and tool-portfolio problem, and it will not be solved by an internal memo about "using AI thoughtfully."
Where Single-Vendor AI Costs a 50โ500 FTE Company Real Money
Homogenization is expensive precisely where differentiation is the product.
Proposals and bids. Your win rate on a competitive RFP is a function of what the buyer cannot get from the other three submissions. If your team and your competitors' teams route drafting through the same assistant, working from the same publicly-scraped conventions, the surviving difference narrows to price. You will observe this as margin pressure and attribute it to the market.
Hiring rubrics and job architecture. Scorecards, competency definitions, and interview guides generated from a shared model converge on a shared idea of a good candidate. Convergent criteria applied across many employers produce correlated rejections โ the same people screened out everywhere, for the same reasons, with no one having made that decision.
Positioning and customer communication. Brand voice is variance from a category baseline. Regression toward that baseline is not a stylistic issue; it is the slow deletion of the thing your marketing spend was buying.
Strategy and option generation. This is the worst place for it. In the option-generation step, variance is the deliverable. A narrower option set is not a smaller cost โ it is a decision made before anyone in the room realized a decision was being made.
Notice the shared property: in every case, per-person productivity metrics improve while the organizational asset degrades. The dashboards will look excellent throughout.
The hidden-loss structure, in one arithmetic
Say drafting time on a proposal falls 30% and your team ships four more bids a quarter. That gain is legible, attributable, and celebrated in the QBR.
Now say win rate on competitive bids drifts from 24% to 21% over three quarters. Nobody attributes that to the drafting tool, because there is no plausible causal story in the room โ the proposals are better written than they were last year, and everyone can see it. The decline gets assigned to pricing pressure, to a competitor's new pricing model, to market softness. Each of those explanations is available, defensible, and unfalsifiable on your data.
That asymmetry is the whole risk. The gain arrives with a clean attribution chain. The loss arrives without one.
Where This Argument Could Be Wrong
Three limits, stated before you act on any of it.
The evidence base is weighted toward creative and ideation tasks โ story writing, brainstorming, design work. That is where homogenization is easiest to measure, not necessarily where it is largest. Whether the same effect size holds for a quarterly close checklist or a support macro is untested, and I would expect it to be smaller and to matter less where standardization is the goal.
A pooled d of 0.334 also conceals heterogeneity across 61 effect sizes. The pooled estimate tells you the direction is reliable. It does not tell you what your team's number is.
Most importantly, the single-vendor framing is my extension of the evidence, not a finding inside it. These studies compare AI-assisted to unassisted work; they do not run a head-to-head of one-model versus multi-model organizations. The inference โ that concentration amplifies convergence โ follows from the anchoring mechanism and from Sourati and Dehghani's argument about shared models, but it is an inference. Treat it as a testable hypothesis about your own output, not a settled result.
None of that changes the conclusion. Convergence is real, it is invisible to your current instrumentation, and it is measurable in your own environment for roughly the cost of an afternoon.
The Portfolio Decision: What to Audit Before the Next Renewal
Four steps. None require budget, and none require anyone to use AI less.
1. Measure your concentration. What share of externally-facing drafting โ proposals, JDs, customer comms, strategy memos โ passes through a single assistant? Most mid-market ops leaders have never asked, and the answer is usually above 80% because procurement standardized on one seat license. That number is your exposure.
2. Fix the order of operations, not the tool. Anchoring runs on sequence. For any task where variance is the point, require independent human generation first, then AI extension and critique. This is a two-line change to a process document and it neutralizes most of the mechanism, because the model can no longer supply the starting point.
3. Diversify inputs deliberately. Different models, and โ more cheaply โ different prompts. A shared "master prompt" circulated as a productivity best practice is a homogenization engine wearing a helpful hat. If your team has one, that is the highest-leverage thing on this list to retire.
4. Instrument spread, not just volume. Pull 20 recent outputs of the same type. Score pairwise similarity โ embeddings if you have the capability, structured human review if you don't โ and compare against a 2024 sample from the same process. You now have a baseline. Repeat quarterly. This is the only step that converts the argument into evidence about your company.
The uncomfortable version
If your 2026 AI plan measures adoption, seat utilization, and hours saved, it is instrumented entirely on the side of the ledger where the gains sit. That is not an oversight anyone will catch, because the losses do not have a metric attached and never generate a complaint.
One Decision This Quarter
Take the single document type where your differentiation matters most โ for most operations leaders that is the proposal. Pull the last twenty. Read them for how much they resemble each other, not for how good each one is.
If they have converged, you did not lose quality. You lost distinctiveness, and you were paying for distinctiveness.
Single-vendor AI makes every one of your people measurably sharper. The question no dashboard in your company is currently built to answer is whether it is also making them interchangeable โ with each other, and with your competitors' people. Answer it with your own documents before your next renewal, because the vendor has no reason to raise it and your productivity metrics never will.