Scovai Scovai
Organizational Behavior 2026-08-04 1 min read

Six in Ten Managers Use AI to Decide Raises and Layoffs. A New I-O Benchmark Shows It's Weakest Exactly There.

DSL

Dr. Sarah Liu

Six in Ten Managers Use AI to Decide Raises and Layoffs. A New I-O Benchmark Shows It's Weakest Exactly There.

Frontier AI models score between 64% and 82% when an employee-feedback question has a clear, verifiable answer. When the task requires weighing incomplete, emotionally loaded, context-dependent signals into a single defensible takeaway, the best of them fall to 33% (PYX Labs, 2026).

That second category has a name in your company. It is called deciding who gets promoted.

Six in ten managers already use AI to make decisions about their direct reports โ€” raises, promotions, layoffs, terminations (Resume Builder, 2025). The capability curve and the usage curve are pointing in opposite directions, and almost nobody has drawn the line between them.

What the PYX-Voice Benchmark Actually Measured

On 15 July 2026, PYX Labs โ€” a research lab sponsored by Perceptyx โ€” released PYX-Voice, the first benchmark built specifically to evaluate how well frontier AI models understand employee feedback. Seven leading models from OpenAI, Google, Anthropic and xAI were run across 84 real employee-listening tasks and graded against 208 criteria written by industrial-organizational psychologists and organizational behavior specialists (Perceptyx, 2026).

The grading criteria are the whole point. Handing a model a spreadsheet of survey responses is trivial. Defining what a correct reading of that data looks like โ€” what a competent I-O psychologist would conclude and what they would refuse to conclude โ€” is the hard part, and it is the part no general-purpose AI benchmark tests.

The headline finding was not that the models are bad. They are not. Reliability declined sharply and predictably as the interpretive load of the task increased (PYX Labs, 2026). The strongest model in the field cleared 76% overall. Every model tested โ€” without exception โ€” recorded synthesis as its single weakest capability.

That consistency matters more than the absolute scores. A weakness shared by every model from four different labs is not a vendor problem you can procure your way out of. It is a property of the task.

The Cliff Sits Exactly Where the Stakes Do

The benchmark's performance profile maps almost perfectly onto a distinction most operations leaders already make intuitively but rarely enforce in tooling.

Models performed strongest where feedback sorts cleanly into well-defined themes โ€” performance enablement questions about goals, tools, metrics, resourcing. These are categorization problems wearing a human-resources costume. There is a right answer, the answer is verifiable, and the model finds it.

Performance degraded on broad, nuanced themes โ€” change, innovation, the diffuse category of comment that says something is wrong without saying what. Here the model must hold several partial, contradictory signals at once and produce a judgment. Scores dropped as low as 33% (PYX Labs, 2026).

The benchmark's own description of the failure is worth quoting, because it names the mechanism rather than the symptom: "The breakdown happens specifically when they have to weigh incomplete, emotional, or context-dependent signals and resolve them into one clear takeaway" (HR Dive, 2026).

Read that description again and ask what a promotion decision is. It is incomplete information, emotionally weighted, heavily context-dependent, resolved into one clear takeaway. The benchmark did not test performance management. It tested the exact cognitive operation performance management consists of.

A 33% score on a judgment task is not a slightly weaker assistant. It is a coin flip with a confident voice.

And the failure mode is worse than the score suggests, because the output looks identical either way. A model synthesizing employee feedback badly does not return an error or a low-confidence flag. It returns a fluent, well-structured, entirely plausible paragraph about team morale. The reviewing manager has no signal distinguishing the 82% case from the 33% case. PYX Labs also recorded rare but meaningful instances of models producing fabricated statistical outputs, or failing to hold to the constraints of the underlying dataset โ€” a figure in the summary that was never in the data (HR Dive, 2026).

Fluency is not calibration. Nothing about the presentation layer degrades when the reasoning does.

Where AI People Decisions Are Already Being Made

Now put that capability profile next to actual manager behavior.

Resume Builder surveyed 1,342 US managers with direct reports in late June 2025. Sixty percent use AI tools to make decisions about their people. Among those who do: 78% use it to determine raises, 77% for promotions, 66% for layoffs, and 64% for terminations (Resume Builder, 2025).

Three further numbers from the same survey define the governance problem precisely.

More than one in five managers say they frequently let AI make the final decision with no human input. Two-thirds of managers using AI to manage people have received no formal training in it. And only 32% report any formal training on using these tools ethically (HR Dive, 2025).

The confidence is the risk, not the usage

More than seven in ten of the managers using AI to help manage their teams expressed confidence that the technology makes fair and unbiased decisions about employees (HR Dive, 2025).

Seventy-plus percent confident. Thirty-three percent capable at the interpretive end. That gap is the entire exposure, and it is not closed by better models โ€” it is closed by better scoping.

Note what this is not. It is not an argument that managers are reckless or that AI has no place in performance management. The managers delegating theme-extraction across 400 open-text survey responses are making a good call; that is the 82% task, and doing it by hand is a waste of a senior person's week. The problem is that the same interface, the same prompt box, and the same institutional permission cover both tasks โ€” and only one of them is safe.

The Case Against Overreading This

Three caveats, stated plainly, because the reflex to discount both sources is partly earned.

PYX Labs is sponsored by Perceptyx, which sells employee-listening technology. A benchmark concluding that employee-feedback interpretation is harder than it looks โ€” and requires I-O expertise โ€” is commercially convenient for a company that sells I-O expertise. Note it, then weigh it against the fact that the benchmark's methodology, task count, and grading criteria are disclosed and the results were published in full, including where models did well.

Resume Builder is a resume-services company, and its study is a self-reported online survey. "Sixty percent of managers use AI for people decisions" measures what managers say they do, in a survey where saying yes carries no cost. Treat the direction as solid and the decimal points as soft.

And PYX-Voice is one benchmark, released weeks ago, not yet subjected to independent replication. A single benchmark is evidence, not proof.

What survives all three caveats is the shape of the finding: capability falls as interpretive load rises, uniformly across labs, while usage is highest exactly where interpretive load peaks. No commercial incentive manufactured that inversion. It would take a strange kind of luck for the sponsor bias and the survey bias to happen to point the same way.

What This Means in a 50โ€“500 FTE Company

At enterprise scale, this problem gets absorbed by process. There is a compensation committee, a calibration session, an HR business partner who reads the same feedback independently, and a legal function that will ask how a termination decision was reached.

At 200 people, there is a manager, a spreadsheet, and a deadline.

That is not a criticism of mid-market operations โ€” it is the structural reality that makes this specific finding more urgent here than at a 20,000-person company. The mid-market has the same AI access, a thinner review layer, and materially less tolerance for a wrongful-termination claim. The tool arrived; the calibration process that would catch its errors did not.

There is also a compounding effect worth naming. A model that under-reads nuanced feedback does not err randomly. It systematically favors the legible over the significant โ€” the clearly written complaint over the hedged one, the articulate employee over the reticent one, the theme with a keyword over the theme without. Run that quietly across two promotion cycles and you have not made a mistake. You have installed a bias with a documentation trail.

Draw the Line This Quarter

Split the work by verifiability, not by tool

The usable boundary is not "AI for HR, yes or no." It is: can the output be checked against the source in under a minute? Clustering 300 comments into themes โ€” checkable, delegate it. Ranking two candidates for one promotion slot on the strength of their feedback โ€” not checkable, keep it human. Write that test down and put it in the manager guidance, because right now two-thirds of your managers using these tools have received no formal training at all (Resume Builder, 2025).

Name the four decisions AI may not close

Raises, promotions, layoffs, terminations. For each, require a documented human rationale that does not cite an AI summary as its evidence. This costs nothing and it directly addresses the one-in-five who currently let the model decide unsupervised.

Audit one cycle backwards

Take the last performance or compensation cycle. Ask each manager one question: where did AI touch this, and what did you check? You are not looking for violations. You are looking for the honest map of where this already happens โ€” which, in most mid-market companies I have looked at, is further along than the leadership team assumes.

One Decision This Quarter

Pick the highest-stakes people decision your managers made in the last ninety days, and ask whether an AI-generated summary sat anywhere in its evidence chain. If it did, ask the second question: could anyone have told, from the output alone, whether that summary was the 82% kind or the 33% kind?

If the answer is no, you do not have an AI problem. You have an unscoped delegation โ€” and the fix is a line in a policy document, not a better model.

Six in ten managers have already crossed that line. Almost none of them were told where it was.

Ready to go beyond the CV?

Scovai's AI-powered Talent Passport reveals what resumes can't: personality, potential, and true job fit.