Sixty-three percent of active job seekers have now sat an AI interview โ up 13 percentage points in six months. Seventy percent were never clearly told that AI would be evaluating them, and 21 percent found out only once the interview had started (Greenhouse, 2026).
Those two numbers describe most mid-market screening funnels. A new paper in Information Systems Research explains what the second one costs you.
Candidates who cannot see how they are being judged embellish. The AI score cannot tell that they did. And the human reviewer you added as a safeguard looks at the score and stops using their own judgment.
That is not three problems. It is one failure mode with three stages, and the middle stage is invisible from your dashboard.
The Transparency Paradox, Tested Across Eight Studies
Akshat Lakhiwal (University of Georgia), Che-Wei Liu (Arizona State), Hillol Bala (Indiana) and Hung-Yue Suen (National Taiwan Normal University) published "From Opacity to Transparency: User Behavior and Downstream Effects in Algorithmic Evaluation" in Information Systems Research on 30 July 2026. Eight studies, including two main experiments and a quasi-field replication, all set in one-way asynchronous video interviews (Lakhiwal et al., Information Systems Research, 2026).
The starting tension is a real one, and it is the reason most employers have stayed quiet about their tooling. Disclose what the algorithm measures and you may reduce candidate anxiety โ or you may simply hand over the test answers. The authors call this the transparency paradox and then go and measure which side wins.
Experiment I: algorithmic evaluation, compared against human evaluation, increased interviewees' stress and deceptive impression management. Transparency counteracted both effects, "aligning them closely with human evaluation."
The disclosure that achieved this was not a technical briefing. Candidates were told the system would look at facial expressions, verbal sentiment and specific keywords, and that it would rate them on teamwork, job-related abilities, work style and personality. That group then displayed the same level of authentic behaviour as candidates who believed a human was reviewing them (UGA Today, 2026).
Lakhiwal's framing of why opacity backfires is worth keeping:
"If you're applying for a job at your dream company and your dream company wants to interview you using AI, you really don't have a lot of choice. You don't really want to say no to it, but you also don't know how it works."
Candidates fill that gap with guesswork. "A lot of embellishment was happening," Lakhiwal said. "They seemed to be throwing the kitchen sink at the situation to try to give the 'evaluator' what it was looking for."
The gaming objection turns out to be backwards. Withholding the criteria did not prevent gaming. It produced gaming, because a candidate optimising against an imagined rubric overshoots in every direction at once.
Your AI Interview Score Couldn't Tell Honest From Embellished
Here is the part that should change how you read your screening reports.
The industry-favoured AI agent used in the study did not penalise applicants who stretched the truth. It scored them as well as candidates who earnestly described having the same qualifications (UGA Today, 2026).
When human evaluators watched the same videos, they could discern it. They generally penalised the embellishers and gave their highest ratings to the candidates who behaved most authentically.
So the human judgment works. That is the encouraging half of the finding, and it survives only until you tell the human what the machine thought.
Then the Humans Stopped Looking
Experiment II moved downstream, comparing three decision configurations: human-only, AI-augmented โ human decision makers assisted by AI scores โ and fully automated.
Transparency kept doing its job on the candidate side across all three. On the evaluator side it ran out of road. In the authors' words, its corrective effect on interview performance "was constrained on the evaluation side." When AI scores were present, evaluators anchored on them, discounting their own judgment, even as the AI scores failed to distinguish honest from deceptive behaviors (Lakhiwal et al., Information Systems Research, 2026).
Read the configurations in order of how common they are in the mid-market. Fully automated screening is rare and widely understood to be risky. Human-only screening is what you are trying to move away from on cost. AI-augmented โ a score, then a human reviewer who signs off โ is what almost everyone has actually bought, and it is the configuration in which a human who can detect embellishment reliably stops doing so.
The safeguard is not neutral. It launders the score.
The Same Failure Showed Up in a Field Experiment
This is not a one-paper result. Lane, Boussioux, Ayoubi and colleagues ran a field experiment with 228 evaluators screening 48 real submissions, comparing human-only evaluation, a black-box model recommendation, and the identical recommendation wrapped in a narrative explanation (Lane et al., HBS Working Paper 25-001).
Black-box recommendations improved decision quality. The same recommendations with an explanation attached did not โ while producing higher compliance. Evaluators disproportionately followed rejection recommendations, substantially increasing false negatives, with the effect strongest in borderline cases.
Two different research teams, two different domains, one mechanism: a fluent machine output substitutes for the human verification step rather than informing it. And false negatives are the one error class a funnel structurally cannot see. A bad hire shows up in six months. A rejected good candidate never shows up at all.
What Candidates Are Asking For Is Not What Protects You
Now put the candidate-side survey data back next to the experiments, because the overlap is uncomfortable.
Thirty-eight percent of candidates say they want to know that a human reviews the AI's evaluation before any decision is made. Thirty-nine percent want a clear explanation of what the AI is measuring (Greenhouse, 2026).
The second request is the one with evidence behind it. Disclosure measurably restored authentic behaviour.
The first request โ human review of the AI's evaluation โ describes, almost exactly, the AI-augmented configuration in which the human defers. It is the safeguard candidates want, the safeguard your vendor will happily sell you, and the safeguard the experiment found wanting. Not because human review is worthless, but because the order matters: reviewing an evaluation is a different cognitive task from evaluating a recording.
Meanwhile 38 percent of candidates have already abandoned a hiring process because it included an AI interview, and the single largest trigger is a pre-recorded video scored by AI with no human present (33 percent). You are losing real candidates to a configuration that also does not work.
The Honest Counter
Four limits, stated plainly.
No effect sizes here, on purpose. The full text sits behind a paywall and the published abstract and press release report directions, not magnitudes โ "a considerable increase," "hundreds of online job seekers." Anyone quoting you a percentage from this study is inventing it. The direction and ordering of the effects is what is established; treat the size as unknown.
One vendor agent, one moment in time. The scoring failure was observed in the specific industry-favoured agent the researchers used. A different tool may discriminate deception better. None of them currently advertise that they can, which is itself informative, but do not generalise a capability claim from a single system.
Prior research supports the gaming objection. Lakhiwal acknowledges it directly: there is existing work showing that telling candidates how they are assessed lets them game the assessment. This paper finds the opposite net effect in this context. One well-designed study does not retire a literature.
Lane et al. is a working paper on adjacent terrain. It evaluates innovation submissions, not job candidates, and the manuscript predates its 2026 circulation. It corroborates the anchoring mechanism; it is not a replication in hiring.
What to Change Before Your Next Screening Round
Four moves. None needs a new budget line, and two are sequencing changes you can make this week.
- Publish what the system measures, in plain language. Not the model name, not the vendor architecture โ Lakhiwal is explicit that candidates do not need that. Tell them which behaviours are observed and which attributes are rated. The study's disclosure condition was four lines long and it restored authentic behaviour. You also fix the 70 percent non-disclosure problem that is driving candidates out of your funnel.
- Make the human rating come first. Have your reviewer score the recording before the AI score is visible, and store both. This is the one change that directly addresses the anchoring result, and it costs nothing but interface order. Note it is Scovai's inference from the finding, not a condition the researchers tested โ the paper shows the anchoring, not the cure.
- Stop treating "a human signed off" as an audit trail. If the human saw the score first, their sign-off carries the score's errors with a person's name attached. What makes a screening decision defensible is an independently recorded human judgment, logged per hire โ Scovai's enterprise posture, including the EU AI Act Article 14 human decision log, documents one way that is structured.
- Audit for false negatives, not accuracy. Pull a sample of candidates your AI screen rejected at the margin and have someone rate the recordings blind. The Lane result says reject-side errors are where compliance concentrates, and your metrics are structurally blind to them.
Your AI interview stack probably has a human in it already. The research question is not whether they are there โ it is whether they formed a judgment before the machine told them what to think.
Go and look at your reviewer's screen. If the score loads above the recording, you do not have a human in the loop. You have a witness.