Scovai Scovai
AI & Operations 2026-09-16 1 min read

Your Extraverts Are the Worst Fit for a Conscientious AI: MIT's New PNAS Experiment Says the Assistant Persona You Rolled Out to Everyone Is a Performance Variable

DSL

Dr. Sarah Liu

Your Extraverts Are the Worst Fit for a Conscientious AI: MIT's New PNAS Experiment Says the Assistant Persona You Rolled Out to Everyone Is a Performance Variable

Pair 1,258 people with AI agents, have them produce 7,266 display ads, then have 1,168 independent raters score the output. The single worst-performing combination in the whole design was not a weak model or an untrained user. It was extraverted humans working with a conscientious AI (Ju & Aral, PNAS, 2026).

Conscientious. The trait every vendor tunes toward by default โ€” careful, thorough, on-task. Given to the wrong person, it produced the lowest-quality work in the experiment.

That is the finding mid-market operations teams should sit with, because almost every one of them has made the opposite bet: one assistant persona, configured once, rolled out to everybody. This study is the first large-scale causal evidence that the persona itself is a performance variable โ€” and that leaving it at the vendor default is a decision, not a neutral state.

What the Experiment Actually Measured

Harang Ju (Johns Hopkins Carey Business School) and Sinan Aral (MIT Sloan, director of the MIT Initiative on the Digital Economy) ran a preregistered randomized experiment: 1,258 participants were measured on the Big Five personality traits, then randomly paired with AI agents prompted to exhibit independently high or low levels of each of those same five traits (Ju & Aral, PNAS, 2026).

The teams had 40 minutes, real-time chat, and synchronized text and image editing to build display ads for a real think tank's year-end report. Output: 7,266 ads. Those ads were then scored on text and image quality by 1,168 independent raters (MIT Sloan / EurekAlert, 2026).

Then the part that makes this more than a lab result. The researchers ran the ads on X for two weeks โ€” nearly five million impressions โ€” and measured click-through rate and cost-per-click. Rated quality predicted market performance: higher text quality lifted click-through by roughly 7 percent and cut cost-per-click by about 30 cents (MIT Sloan / EurekAlert, 2026).

Two design choices deserve emphasis, because they are what separate this from the persona commentary already in circulation. The trait levels were randomly assigned and varied independently, so the comparison is not between self-selected users of different tools โ€” it is causal. And the study measured the humans on the same Big Five instrument used to prompt the agents, which is what makes the interaction estimable at all. Most vendor research reports how users feel about a persona. This one reports what the pairing did to the work.

So the causal chain is complete end to end. Persona pairing moved quality. Quality moved money.

The Pairing Table: Which Combinations Helped, Which Hurt

The results do not reduce to "match like with like," which is what most people assume before reading them.

The strongest pairings were extraverted humans with extraverted AI, conscientious humans with conscientious AI, and open humans with conscientious AI โ€” that last one a complement, not a mirror (MIT Sloan / EurekAlert, 2026).

The pairings that significantly degraded quality, in the order the paper gives them:

  1. Extraverted humans + conscientious AI โ€” the lowest-quality output in the study
  2. Conscientious humans + agreeable AI
  3. Neurotic humans + conscientious AI

(Ju & Aral, PNAS, 2026)

Read the Second Row Twice

Conscientious AI appears in two of the three worst pairings and in two of the three best. There is no globally good persona in this data. A conscientious agent given to a conscientious person was among the best combinations available; the same agent given to an extravert was the worst. Identical configuration, opposite result, determined entirely by who received it.

And notice what the second-worst row does to the standard vendor answer. Agreeable AI โ€” helpful, accommodating, never pushes back โ€” degraded the output of conscientious people. The persona most tuned for user satisfaction was actively unhelpful to the population most likely to be doing your careful work.

Why One Org-Wide Persona Is a Silent Performance Tax

Here is the operational translation, and it is uncomfortable precisely because it costs nothing to fix and nobody is measuring it.

When you deploy a single assistant configuration across 300 people, you are not running a neutral rollout. You are running an unrandomized experiment in which some fraction of your workforce has been handed the pairing that this study found produces the worst output โ€” and neither they nor you have any instrument that would detect it. The failure shows up as slightly weaker drafts, slightly more rework, a vague sense from certain teams that the tool "isn't that useful." It never shows up as a line item.

That rework lands in the most expensive place. McKinsey QuantumBlack's cost breakdown of agentic banking workflows put token spend at 20 to 25 percent of an agent's variable run cost, and human oversight by functional and risk experts at 70 to 75 percent (McKinsey QuantumBlack, 2026). Output quality is not a soft metric in that structure. Every point of quality you lose at generation is paid back at review rates, by the senior person doing the checking.

And the checking is not going away. In MIT Technology Review Insights and Microsoft's study of 300 technology executives ranking 101 agentic tasks, 59 percent already plan permanent human oversight, with confidence tracking task verifiability rather than raw model capability (MIT Technology Review Insights, 2026).

A persona mismatch is therefore not a user-experience complaint. It is a quiet multiplier on the single largest cost line in your AI stack.

The Variance Is Already Inside the Job Title

If you want evidence that this is measurable in your own company rather than only in a lab, it already exists in the adoption data. Bick, Blandin, Deming and Schumacher, linking a nationally representative survey to detailed O*NET tasks, found that generative-AI exposure measures explain only about half of the variation in adoption across workers โ€” people doing the same job, with the same tools, adopt systematically differently (NBER Working Paper 35677, 2026).

Half the variance sits inside the job title, not between job titles. Most ops teams file that under motivation or training. The PNAS result offers a less flattering and more actionable explanation for part of it: the tool is configured for a person who is not the one using it.

That reframe changes what you do with a low-adoption cluster. Sending another enablement session to a team that is quietly getting worse output from the default agent does not address the cause โ€” it adds hours to people whose drafts will still come back needing more rework than their colleagues'.

The Honest Counter

Three limits, and they matter.

The task was ad creation, not operations. Forty-minute creative sessions producing display ads are not your quarterly close, your vendor onboarding, or your ticket triage. That personality pairing moved quality in a creative task is a finding; that it moves quality in your workflows is an inference. Treat it as a strong prior worth testing, not a settled result.

The authors have a commercial position. Ju and Aral cofounded Pairium, which builds and distributes the Pairit experimentation platform used in the study (Ju & Aral, PNAS, 2026). The work is preregistered and peer-reviewed at PNAS, which constrains the obvious failure modes โ€” but the researchers with the finding also sell the remedy, and that belongs in your reading of it.

Matching everyone perfectly has its own cost. Work in the same literature finds that diverse AI personas mitigate the homogenization effect in humanโ€“AI collaborative ideation (Wan & Kalman, 2026). Optimize every pairing for individual output quality and you may compress the variance across your teams' thinking โ€” better drafts, more similar drafts. If your function needs range rather than polish, that trade is real.

None of which rescues the status quo. The default configuration is not a hedge against these limits; it is the one option guaranteed to be wrong for part of your workforce, with no data telling you which part.

What to Decide This Quarter

Four moves, none requiring new vendor spend.

  1. Stop treating the persona as IT configuration. The assistant's system prompt is a performance parameter with measured effects on output quality. Whoever owns AI enablement owns it โ€” and owns justifying the current setting.
  2. Run one pairing test on a function that already produces gradeable output. Content, support responses, first-draft analysis. Two personas, randomly assigned, one quarter, blind quality ratings. You are replicating a published design at small scale, which is the cheapest research you will ever commission.
  3. Look hardest at your extraverts and your high-neuroticism staff in high-output roles. Those are the two human profiles the study flagged against the default-conscientious agent. If a visibly capable person is quietly getting little from the tool, mismatch is now a live hypothesis rather than a training problem.
  4. Offer persona choice before you build persona assignment. Full trait-matched allocation needs personality data most mid-market firms do not hold and may not want to. Two or three selectable assistant profiles capture part of the effect with none of the governance exposure.

AI personality pairing is now a measured variable with a price attached. You are already setting it โ€” once, globally, by accepting whatever the vendor shipped.

The question for this quarter is not whether the persona matters. PNAS has answered that. It is whether you are going to keep making that call by default, for everyone, and calling the result the model's fault.

Ready to go beyond the CV?

Scovai's AI-powered Talent Passport reveals what resumes can't: personality, potential, and true job fit.