The next button
We scored 8,428 AI conversations at work. Not the AI, the humans. Half of our most "agentic" users turned out to be a next button. Science has a name for it, twenty years of receipts, and one fix that everybody hates.

We pulled our company's complete claude.ai enterprise export. 8,428 conversations, 131 accounts, fifteen months, every message.
Then we did something slightly unfair. We scored the humans.
An LLM judge read only the human's side of each conversation. The AI's replies were hidden. Four scores, 0 to 3 each: did you scope the task, did you catch errors, did you check the output, did you supply context the model couldn't know.
The delegation numbers looked great at first. Tools firing, multi-step agent runs, real work moving. Exactly the chart you'd show a board.
Then we read the transcripts behind the chart. In 48% of the delegated conversations, the human did one thing. "Fix it." "Still broken." "Try again." No question. No check. No pushback. A next button.
I thought we'd discovered a new failure mode. We hadn't. Science has been studying it for twenty years.
It has a name
The research literature calls it overreliance: accepting AI output without evaluating it. Microsoft's research group reviewed about 60 studies on it back in 2022, before chat AI was even good.
The findings are brutal. People who accept wrong AI output do worse than people working with no AI at all. And they're often slower, because unwinding a wrong answer costs more than writing a right one from scratch.
A 2024 meta-analysis in Nature Human Behaviour pooled 106 studies. On average, human plus AI performed worse than the better of the two alone. The losses concentrate exactly where companies want AI most: decision tasks.
The next button doesn't just fail to add value. It subtracts it.
Nobody checks
Our weakest score, by far, was verification. Briefing the AI: 2.08 out of 3. We're good at asking. Checking what came back: 1.03 out of 3.
Same gap in every department. Every role. Every site. And flat across six months of data. Nobody learned on their own.
The research says that's exactly what should happen. A CHI 2025 study of 319 knowledge workers found that the more people trust the AI, the less they think critically about its output. Trust doesn't sharpen the checking muscle. It replaces it.
It gets worse. A randomized experiment with 400 people found that after making a wrong reliance decision, people became more confident, not less. Being wrong didn't teach them. It reassured them.
So no, this will not fix itself with experience. Experience is the thing making it worse.
Two skills, not one
Here's the finding that changed how I read our own numbers. Accepting good AI output and catching bad AI output are different skills. In the research they move independently. In our corpus, the correlation between briefing well and checking well is 0.36. Weak.
Your best prompt engineer can be your worst verifier. Ours often is.
Explanations and confidence scores, the things every AI product ships, help with the first skill only. They make people better at accepting good advice. They do nothing for rejecting bad advice. Which is the skill that was missing in the first place.
What actually works
Warnings? Disclaimers? "AI can make mistakes"? Measured, repeatedly: they reduce blind acceptance a little and improve the actual skill not at all.
The one intervention with replicated evidence is what researchers call cognitive forcing functions. Checklists. Mandatory review steps. A pause before you're allowed to accept. Anything that makes you stop and think when every incentive says keep moving.
The studies are honest about the catch: people hate it. It's friction, and it feels like distrust of a tool that's right most of the time.
But "right most of the time" is exactly the problem. A tool that's right 90% of the time trains you to stop checking. The 10% lands in production with your name on the commit.
The mirror
One more number from our corpus. Our ten heaviest users produced 40% of all conversations. Their judgment scores were identical to everyone else's. Volume taught them nothing.
We spent a year asking how good the AI is getting. Wrong question. The model was never the variable.
The next button is us.