Here’s another interesting, faintly dystopian AI term: cognitive surrender. You can probably guess what it means. It comes out of an attempt to extend one of the most familiar frameworks in behavioral science, and I suspect it has travelled as fast as it has because most of us did the thing it describes fairly recently. I have asked a model for an answer, read it, decided it sounded right, and moved on without checking a single part of it.
Steven Shaw and Gideon Nave (The Wharton School) propose Tri‑System Theory, extending dual‑process accounts of reasoning by positing System 3: artificial cognition that operates outside the brain, and that can supplement or supplant internal processes. Its key prediction is cognitive surrender, meaning adopting AI outputs with minimal scrutiny, overriding both intuition (System 1) and deliberation (System 2).
Across three preregistered experiments using an adapted Cognitive Reflection Test (N = 1,372; 9,593 trials), the authors randomized AI accuracy using hidden seed prompts. Participants chose to consult the AI assistant on a majority of trials. Against a no‑AI baseline, accuracy rose 25 percentage points when the AI was accurate and fell 15 points when it erred, which the authors treat as the behavioral signature of cognitive surrender (Cohen’s h = 0.81). Engaging System 3 also raised participants’ confidence, including after errors.
Time pressure (Study 2) and per‑item incentives with feedback (Study 3) shifted baseline performance without eliminating the pattern. Accurate AI buffered the costs of time pressure and amplified incentive gains. Faulty AI reduced accuracy regardless of the situational moderators. Participants with higher trust in AI and lower need for cognition and fluid intelligence surrendered more.
This summary was generated with AI on August 13, 2026, and checked against the source paper by our editorial team.
The part I’m not sold on
I’m not 100% convinced this third system is fully new. We’ve always had ways of delegating our thinking, and by the paper’s own definition, external reasoning that supplements internal processes covers some very old technology.
Still, at some point AI stops being a tool and becomes a self‑contained thinky thing, closer to collaborating with a person than looking something up. The behavioral signature Shaw and Nave land on is the part I would keep, because it turns cognitive surrender from a vibe into a number I can argue with.
Whether updating dual‑process theory is a legitimate advance or mostly for pizazz, something is clearly shifting.
What the number actually shows
Participants worked through an adapted Cognitive Reflection Test, the classic instrument for catching the moment someone overrides a tempting wrong answer and thinks it through. One group had no AI access, while another could consult an assistant as often as they liked, without knowing that its accuracy had been randomized behind the scenes.
Participants’ answers followed the assistant in whichever direction it went. Accuracy climbed 25 points above baseline when the AI was right and dropped 15 points below it when the AI was wrong, a spread of roughly 40 points driven entirely by something the participants could not see. The effect size is large, and I think that single contrast carries most of the paper’s argument.
The confidence result unsettles me more. Consulting the assistant raised confidence even when it had led people to the wrong answer, which means the failure never announces itself. Someone ends up more sure of a worse answer, and that is exactly the state in which I would never think to go back and check.
behavior change 101
Start your behavior change journey at the right place
Accuracy rose 25 points when the AI was right and fell 15 when it was wrong, and confidence went up either way.
I would point managers at Studies 2 and 3, because they test the two fixes people reach for first. Under time pressure, and then with per‑item incentives and feedback, baseline performance moved about as I would have expected, but the pattern did not break. An accurate assistant cushioned the cost of time pressure and made the incentives pay off more, while a faulty one reduced accuracy regardless. Paying people to be careful and telling them how they did was not enough to stop them following a confident wrong answer.
I expect the moderator findings to get misused. Higher trust in AI, lower need for cognition, and lower fluid intelligence all predicted more surrender, which reads easily as a story about other people being gullible. I think it more plausibly describes a disposition to check that varies quite a lot between people, and that nobody screens for at the moment someone is handed a tool.
The paper was posted to SSRN on February 2, 2026 (dated January 11, 2026, last revised February 10, 2026) as a Wharton School research paper. At the time this article was written, on August 13, 2026, it remains a preprint and has not completed peer review. Findings described here should be read with that in mind.
How people are reacting
When I put the paper in front of people, the replies came back unusually productive, because most of those arguing with the label ended up sharpening the underlying idea for me. Four threads did most of the work.
The reading I heard most often treats System 3 as something happening inside System 2 rather than beside it. A design researcher at a brand strategy consultancy described it to me as externalised, scaffolded System 2, reshaping the deliberative process without forming a faculty of its own.
He makes a fair historical point too. Cognitive augmentation under low friction is how humans have always thought. We offload memory to marks, knots, ledgers and writing, and we offload navigation to maps, instruments and GPS. He ends on responsibility, since something has to remain answerable for the result, and in his view that something is a person, whatever the tooling does.
I would note that this reading strengthens the paper’s empirical claim while weakening its theoretical one, since the 40‑point swing happened either way and only the numbering goes away.
A product director at a software company put the mechanism to me plainly: when the output comes back fully formed, it bypasses the friction that usually forces System 2 to wake up. Delegation is old, and what feels new to him is delegation this smooth.
That fits the confidence result closely, since the assistant in these studies delivered wrong answers in the same assured register as right ones. It also puts part of the problem in the interface, which I find easier to change than anyone’s judgment.
A governance lead at a large consultancy made the institutional version of the argument to me, and noted that it predates AI by decades. Committees have always deferred to the most experienced person in the room, and when they do, the reasoning lives in one head and walks out of the building with it. AI delegation differs because it is invisible at the seams. With a person, you at least know who decided, whereas here the behavioral signature is the only evidence that anything was surrendered at all.
His conclusion was the most concrete thing anyone sent me. If you cannot tell afterwards which decisions were reasoned and which were adopted, you have a provenance problem that no AI policy will solve. He wants a record of which system was active, what criteria were applied, and why.
A policy advisor at an AI governance nonprofit framed the same gap as authority and consequence coming apart. People carry the cost of errors made by systems they cannot audit, which raises a harder question about why so many workflows now require an opaque tool to get anything done. A clinical psychologist offered the individual version, calling it learned helplessness, and comparing it to the way an overbearing caregiver removes the mental load a person needs in order to assert any authority of their own.
A smaller set of replies questions the base rather than the addition. A clinician‑researcher at a teaching hospital called dual‑process models a relic and was unwilling to accept a third system when he does not accept the first two as described, extending the same doubt to working and long‑term memory as separable compartments. A consultant at a systems‑safety firm took issue with the idea of AI scaffolding existing systems, arguing that once the goal is to support a conclusion already reached, the result is automated justification.
Both critiques take aim at the theory, and both leave the measurement standing. Participants who followed a confident wrong answer still ended up 15 points below where they would have landed alone, and organizations have to manage that however the systems end up being counted.
Across everything people sent back, I keep noticing that the disagreement is really about where the phenomenon belongs, whether that is a theory of mind, an interface problem, or a governance question. I think it sits in all of those places at once, which is why it resists a clean policy fix.
Will human oversight fall short?
Most new AI policies work by having a person in the loop to decrease the risk that AI systems will go rogue on us (take the EU AI Act, for example). But these types of regulations can’t ensure that the humans overseeing AI won’t “surrender” to its outputs, especially when time pressures, fatigue, and extended use come into play.
Study 2 is the part that makes me uncomfortable, since time pressure did not break the pattern and an accurate assistant actually cushioned its costs. A reviewer working against the clock with a mostly reliable system will look like a functioning safeguard right up until the output is wrong, and the paperwork will record a human sign‑off either way.
Obviously, this isn’t an argument to throw caution to the wind and give up on humans in the loop entirely (that would be a complete surrender, not just a cognitive one). If anything, it’s a reminder that making sure AI does what we want it to do can’t solely be technical.



















