Two Wharton researchers want to add a third system to dual-process theory. You can argue with the label, though the 40-percentage-point accuracy gap between people given a right AI and people given a wrong one is harder to wave away.
People have always reasoned with help from outside their own heads. Arithmetic moved from the abacus to the calculator, and looking things up moved from the library to the search engine. Each tool got better at its one job, while the reasoning itself stayed with the person using it. Tools that can hand back a finished line of reasoning change that arrangement, and it was probably only a matter of time before one could produce better reasoning than its user on some tasks.
For organizations that count on a person signing off on AI outputs, a new paper suggests that sign-off may involve less checking than it appears to. Steven Shaw and Gideon Nave (The Wharton School) call the underlying behavior cognitive surrender, meaning accepting an AI's answer with little or no scrutiny. The term comes out of an attempt to extend one of the most familiar frameworks in behavioral science. It has probably travelled as fast as it has because most people who use these tools have done the thing it describes: asked a model for an answer, read it, decided it sounded right, and moved on without checking any part of it.
Shaw and Nave propose Tri-System Theory. Dual-process theory describes two kinds of thinking, System 1 (fast and intuitive) and System 2 (slow and deliberate). The authors add System 3: artificial reasoning that happens outside the brain and can supplement or replace the other two. Its key prediction is cognitive surrender, meaning adopting AI outputs with little scrutiny and overriding both intuition and deliberation.
The authors ran three experiments with their analysis plans fixed in advance. About 1,400 people answered close to 9,600 questions from an adapted Cognitive Reflection Test, a set of puzzles designed so that the first answer that comes to mind is wrong. The researchers secretly controlled whether an AI assistant gave right or wrong answers, and participants chose to consult it on most questions. Compared with people working without AI, accuracy rose 25 percentage points when the AI was right and fell 15 points when it was wrong. The authors classify this as a large effect and treat it as the behavioral signature of cognitive surrender. Using the AI also raised participants' confidence, including after errors.
Time pressure (Study 2) and per-question rewards with feedback (Study 3) shifted performance without eliminating the pattern. An accurate AI softened the cost of time pressure and increased the gains from rewards. A faulty AI reduced accuracy in every condition. People who trusted AI more surrendered more. So did people who scored lower on two measures: need for cognition, meaning how much someone enjoys effortful thinking, and fluid intelligence, meaning the ability to reason through unfamiliar problems.
This summary was generated with AI on August 13, 2026, and checked against the source paper by The Decision Lab's editorial team.
Incentives and time pressure did not break the pattern
Managers should look first at Studies 2 and 3, because they test the two fixes organizations usually reach for first. Adding time pressure, and then adding per-question rewards with feedback, moved baseline performance roughly as expected, but neither broke the pattern. An accurate assistant cushioned the cost of time pressure and made the rewards pay off more, while a faulty one reduced accuracy regardless. Paying people to be careful and telling them how they did was not enough to stop them following a confident wrong answer.
The findings on who surrendered most are likely to be misused. People who trusted AI more surrendered more, and so did people lower in need for cognition or fluid intelligence. That reads easily as a story about other people being gullible. The results more plausibly describe a disposition to check that varies a lot between people, and that nobody screens for at the moment someone is handed a tool.
The paper was posted to the Social Science Research Network (SSRN) on February 2, 2026, as a Wharton School research paper. It is dated January 11, 2026, and was last revised February 10, 2026. When this article was written, on August 13, 2026, it was still a preprint and had not completed peer review.
behavior change 101
Start your behavior change journey at the right place
The third system may be older than it looks
The theory holds up less well than the data behind it. By the paper's own definition, outside reasoning that supplements what happens in the head covers some very old technology, and it is not obvious that renumbering the theory adds much beyond pizzazz.
Still, a model that works through a problem and hands back a finished answer is a self-contained thinky thing, closer to collaborating with a person than looking something up. The paper's more lasting contribution is the behavioral signature Shaw and Nave measure, since it gives cognitive surrender a number someone can argue with.
How others are reacting
When the paper was shared, most of the people who argued with the label ended up sharpening the idea underneath it. The replies fell into four threads.
The most common reading places System 3 inside System 2 rather than beside it. A design researcher at a brand strategy consultancy described it as System 2 moved partly outside the head. In his view, it changes how deliberation works without becoming a separate kind of thinking.
He makes a fair historical point too. Humans have always offloaded memory to marks, knots, ledgers and writing, and navigation to maps and GPS. He ends on responsibility: something has to stay answerable for the result, and in his view that something is a person, whatever the tools do.
This reading strengthens the paper's empirical claim while weakening its theoretical one. The 40-point gap happened either way, and only the numbering goes away.
A product director at a software company put the mechanism plainly. When the output comes back fully formed, it bypasses the friction that usually forces System 2 to wake up. Delegation is old, and what feels new to him is delegation this smooth. Psychologists call this processing fluency: information that is easy to take in tends to feel true.
That fits the confidence result closely, since the assistant in these studies delivered wrong answers in the same assured tone as right ones. It also puts part of the problem in the interface, which is easier to change than anyone's judgment.
A governance lead at a large consultancy made the institutional version of the argument, and noted that it predates AI by decades. Committees have always deferred to the most experienced person in the room. When they do, the reasoning lives in one head and walks out of the building with it. AI delegation differs because it is invisible at the seams. With a person, you at least know who decided, whereas with AI the behavioral signature is the only evidence that anything was surrendered at all.
He argues that if you cannot tell afterwards which decisions were reasoned and which were simply adopted, you have a provenance problem that no AI policy will solve. By provenance he means a record of where a decision came from. He wants each decision logged with whether it was reasoned or adopted, along with the criteria behind it.
A policy advisor at an AI governance nonprofit framed the same gap as authority and consequence coming apart. People carry the cost of errors made by systems they cannot audit. That raises a harder question about why so many workflows now require an opaque tool to get anything done.
A clinical psychologist offered the individual version and called it learned helplessness, the state in which someone stops trying to act because past efforts made no difference. The comparison was to an overbearing caregiver who takes on so much of a person's mental load that the person loses the footing needed to assert any authority of their own.
A smaller set of replies went after the foundation. A clinician-researcher at a teaching hospital called dual-process models a relic. He saw no reason to accept a third system when he does not accept the first two as described. He extended the same doubt to the idea of working memory and long-term memory as separate compartments. A consultant at a systems-safety firm objected to the idea of AI supporting existing ways of thinking. In this consultant's view, once the goal is to back up a conclusion already reached, the result is automated justification.
Both critiques take aim at the theory, and both leave the measurement standing. Participants who followed a confident wrong answer still scored 15 points below people working without AI. Organizations have to manage that however the systems end up being counted.
Most of the disagreement was about where cognitive surrender belongs. Some placed it in a theory of the mind, and others in the design and governance of the tools people use. It probably belongs in both, and that split is part of why no single policy fixes it.
A human sign-off can hide the same surrender
Many new AI rules rely on a person in the loop as the main safeguard. Article 14(4)(b) of the EU AI Act requires oversight measures to enable the person to remain aware of the possible tendency to automatically rely or over-rely on the output of a high-risk AI system. This tendency is called automation bias. A rule can require that awareness, but it cannot make a reviewer act on it, and time pressure is an ordinary feature of review work.
The Study 2 result applies directly. A reviewer working against the clock with a mostly reliable system will look like a functioning safeguard right up until the output is wrong. Because confidence rose even after errors, the reviewer is unlikely to notice, and the paperwork will record a human sign-off either way.
Humans should stay in the loop, though a sign-off cannot be read as proof that anyone checked. Shaw and Nave could only see surrender by planting wrong answers that participants did not know about. An organization can run the same test by feeding a small, known share of deliberately wrong outputs into a review queue and counting how many come back approved. That means putting errors on purpose into a process built to catch them, and few compliance teams are likely to be comfortable approving it.



















