From Tool to Thinker: When AI Becomes Part of How We Reason

Two Wharton researchers want to add a third system to dual‑process theory. You can argue with the label, though the 40‑point accuracy swing they measured is harder to wave away.

programmer does ai systems checkup

Here’s another interesting, faintly dystopian AI term: cognitive surrender. You can probably guess what it means. It comes out of an attempt to extend one of the most familiar frameworks in behavioral science, and I suspect it has travelled as fast as it has because most of us did the thing it describes fairly recently. I have asked a model for an answer, read it, decided it sounded right, and moved on without checking a single part of it.

The paper in 200 words

Steven Shaw and Gideon Nave (The Wharton School) propose Tri‑System Theory, extending dual‑process accounts of reasoning by positing System 3: artificial cognition that operates outside the brain, and that can supplement or supplant internal processes. Its key prediction is cognitive surrender, meaning adopting AI outputs with minimal scrutiny, overriding both intuition (System 1) and deliberation (System 2).

Across three preregistered experiments using an adapted Cognitive Reflection Test (N = 1,372; 9,593 trials), the authors randomized AI accuracy using hidden seed prompts. Participants chose to consult the AI assistant on a majority of trials. Against a no‑AI baseline, accuracy rose 25 percentage points when the AI was accurate and fell 15 points when it erred, which the authors treat as the behavioral signature of cognitive surrender (Cohen’s h = 0.81). Engaging System 3 also raised participants’ confidence, including after errors.

Time pressure (Study 2) and per‑item incentives with feedback (Study 3) shifted baseline performance without eliminating the pattern. Accurate AI buffered the costs of time pressure and amplified incentive gains. Faulty AI reduced accuracy regardless of the situational moderators. Participants with higher trust in AI and lower need for cognition and fluid intelligence surrendered more.

This summary was generated with AI on August 13, 2026, and checked against the source paper by our editorial team.

The part I’m not sold on

I’m not 100% convinced this third system is fully new. We’ve always had ways of delegating our thinking, and by the paper’s own definition, external reasoning that supplements internal processes covers some very old technology.

Still, at some point AI stops being a tool and becomes a self‑contained thinky thing, closer to collaborating with a person than looking something up. The behavioral signature Shaw and Nave land on is the part I would keep, because it turns cognitive surrender from a vibe into a number I can argue with.

Whether updating dual‑process theory is a legitimate advance or mostly for pizazz, something is clearly shifting.

What the number actually shows

Participants worked through an adapted Cognitive Reflection Test, the classic instrument for catching the moment someone overrides a tempting wrong answer and thinks it through. One group had no AI access, while another could consult an assistant as often as they liked, without knowing that its accuracy had been randomized behind the scenes.

Participants’ answers followed the assistant in whichever direction it went. Accuracy climbed 25 points above baseline when the AI was right and dropped 15 points below it when the AI was wrong, a spread of roughly 40 points driven entirely by something the participants could not see. The effect size is large, and I think that single contrast carries most of the paper’s argument.

The confidence result unsettles me more. Consulting the assistant raised confidence even when it had led people to the wrong answer, which means the failure never announces itself. Someone ends up more sure of a worse answer, and that is exactly the state in which I would never think to go back and check.

behavior change 101

Start your behavior change journey at the right place

Accuracy rose 25 points when the AI was right and fell 15 when it was wrong, and confidence went up either way.

I would point managers at Studies 2 and 3, because they test the two fixes people reach for first. Under time pressure, and then with per‑item incentives and feedback, baseline performance moved about as I would have expected, but the pattern did not break. An accurate assistant cushioned the cost of time pressure and made the incentives pay off more, while a faulty one reduced accuracy regardless. Paying people to be careful and telling them how they did was not enough to stop them following a confident wrong answer.

I expect the moderator findings to get misused. Higher trust in AI, lower need for cognition, and lower fluid intelligence all predicted more surrender, which reads easily as a story about other people being gullible. I think it more plausibly describes a disposition to check that varies quite a lot between people, and that nobody screens for at the moment someone is handed a tool.

Status Note

The paper was posted to SSRN on February 2, 2026 (dated January 11, 2026, last revised February 10, 2026) as a Wharton School research paper. At the time this article was written, on August 13, 2026, it remains a preprint and has not completed peer review. Findings described here should be read with that in mind.

How people are reacting

When I put the paper in front of people, the replies came back unusually productive, because most of those arguing with the label ended up sharpening the underlying idea for me. Four threads did most of the work.

System 2, outsourced

The reading I heard most often treats System 3 as something happening inside System 2 rather than beside it. A design researcher at a brand strategy consultancy described it to me as externalised, scaffolded System 2, reshaping the deliberative process without forming a faculty of its own.

He makes a fair historical point too. Cognitive augmentation under low friction is how humans have always thought. We offload memory to marks, knots, ledgers and writing, and we offload navigation to maps, instruments and GPS. He ends on responsibility, since something has to remain answerable for the result, and in his view that something is a person, whatever the tooling does.

I would note that this reading strengthens the paper’s empirical claim while weakening its theoretical one, since the 40‑point swing happened either way and only the numbering goes away.

Fluency is doing the work

A product director at a software company put the mechanism to me plainly: when the output comes back fully formed, it bypasses the friction that usually forces System 2 to wake up. Delegation is old, and what feels new to him is delegation this smooth.

That fits the confidence result closely, since the assistant in these studies delivered wrong answers in the same assured register as right ones. It also puts part of the problem in the interface, which I find easier to change than anyone’s judgment.

The provenance problem

A governance lead at a large consultancy made the institutional version of the argument to me, and noted that it predates AI by decades. Committees have always deferred to the most experienced person in the room, and when they do, the reasoning lives in one head and walks out of the building with it. AI delegation differs because it is invisible at the seams. With a person, you at least know who decided, whereas here the behavioral signature is the only evidence that anything was surrendered at all.

His conclusion was the most concrete thing anyone sent me. If you cannot tell afterwards which decisions were reasoned and which were adopted, you have a provenance problem that no AI policy will solve. He wants a record of which system was active, what criteria were applied, and why.

A policy advisor at an AI governance nonprofit framed the same gap as authority and consequence coming apart. People carry the cost of errors made by systems they cannot audit, which raises a harder question about why so many workflows now require an opaque tool to get anything done. A clinical psychologist offered the individual version, calling it learned helplessness, and comparing it to the way an overbearing caregiver removes the mental load a person needs in order to assert any authority of their own.

Whether the foundation holds

A smaller set of replies questions the base rather than the addition. A clinician‑researcher at a teaching hospital called dual‑process models a relic and was unwilling to accept a third system when he does not accept the first two as described, extending the same doubt to working and long‑term memory as separable compartments. A consultant at a systems‑safety firm took issue with the idea of AI scaffolding existing systems, arguing that once the goal is to support a conclusion already reached, the result is automated justification.

Both critiques take aim at the theory, and both leave the measurement standing. Participants who followed a confident wrong answer still ended up 15 points below where they would have landed alone, and organizations have to manage that however the systems end up being counted.

Where this lands

Across everything people sent back, I keep noticing that the disagreement is really about where the phenomenon belongs, whether that is a theory of mind, an interface problem, or a governance question. I think it sits in all of those places at once, which is why it resists a clean policy fix.

Will human oversight fall short?

Most new AI policies work by having a person in the loop to decrease the risk that AI systems will go rogue on us (take the EU AI Act, for example). But these types of regulations can’t ensure that the humans overseeing AI won’t “surrender” to its outputs, especially when time pressures, fatigue, and extended use come into play.

Study 2 is the part that makes me uncomfortable, since time pressure did not break the pattern and an accurate assistant actually cushioned its costs. A reviewer working against the clock with a mostly reliable system will look like a functioning safeguard right up until the output is wrong, and the paperwork will record a human sign‑off either way.

Obviously, this isn’t an argument to throw caution to the wind and give up on humans in the loop entirely (that would be a complete surrender, not just a cognitive one). If anything, it’s a reminder that making sure AI does what we want it to do can’t solely be technical.

References

Shaw, S. D., & Nave, G. (2026). Thinking — fast, slow, and artificial: How AI is reshaping human reasoning and the rise of cognitive surrender (The Wharton School Research Paper). SSRN. https://doi.org/10.2139/ssrn.6097646

About us

We are the leading applied research & innovation consultancy

Our insights are leveraged by the most ambitious organizations

Image

I was blown away with their application and translation of behavioral science into practice. They took a very complex ecosystem and created a series of interventions using an innovative mix of the latest research and creative client co-creation. I was so impressed at the final product they created, which was hugely comprehensive despite the large scope of the client being of the world's most far-reaching and best known consumer brands. I'm excited to see what we can create together in the future.

Heather McKee

BEHAVIORAL SCIENTIST

GLOBAL COFFEEHOUSE CHAIN PROJECT

OUR CLIENT SUCCESS

$0M

Annual Revenue Increase

By launching a behavioral science practice at the core of the organization, we helped one of the largest insurers in North America realize $30M increase in annual revenue.

0%

Increase in Monthly Users

By redesigning North America's first national digital platform for mental health, we achieved a 52% lift in monthly users and an 83% improvement on clinical assessment.

0%

Reduction In Design Time

By designing a new process and getting buy-in from the C-Suite team, we helped one of the largest smartphone manufacturers in the world reduce software design time by 75%.

0%

Reduction in Client Drop-Off

By implementing targeted nudges based on proactive interventions, we reduced drop-off rates for 450,000 clients belonging to USA's oldest debt consolidation organizations by 46%

Read Next

Person with Gloves Planting a Tree
Insight

The Dos and Don’ts of Corporate Social Responsibility

Corporate Social Responsibility has become a buzzword for business executives. To remain competitive in this changing economic culture, firms tout their CSR programs, funnelling money into philanthropy, social movements, advertising their new green policies, and emphasizing their commitment to local communities.

Notes illustration

Eager to learn about how behavioral science can help your organization?