Productivity · 10 min read
Slow Thinking for AI-Assisted Decisions
AI accelerates fast thinking and may suppress deliberate review. Research-backed strategies help you decide when and how to slow down before accepting AI output.

When AI generates an answer in milliseconds, the cognitive pressure to accept it is real. Research suggests that simply having access to an AI output — even an unreliable one — may inflate human confidence. Knowing when to slow down, and how to structurally enforce that pause, may be one of the most important governance decisions teams make as AI becomes embedded in daily work.
The Two-System Model No Longer Tells the Full Story
Daniel Kahneman's framework of fast and slow thinking — the distinction between quick, intuitive judgments and careful, deliberate reasoning — has shaped how economists, managers, and executives understand judgment and choice for decades. System 1 handles the automatic, pattern-matching responses; System 2 kicks in for complex, effortful reasoning. Research estimates that around 96% of our thinking runs through System 1, which means deliberate review is always the exception, not the default.
Kahneman's original model was built for a world in which all thinking happens inside the human mind — a world that no longer exists. New research from Wharton postdoctoral researcher Steven Shaw and professor Gideon Nave argues the framework is missing something fundamental. Shaw and Nave propose a Tri-System Theory: engaging with AI has created a third mode of human cognition — System 3 — that influences how both intuition and deliberation operate, and that is now as influential in today's decision-making as the original two systems.
System 3 doesn't simply add speed. It changes the character of both fast and slow thinking by introducing an external cognitive agent whose outputs carry an implicit authority that human outputs do not. Understanding this shift is the starting point for any serious conversation about AI governance and workflow design.
Cognitive Surrender: Confidence Without Accuracy
The most unsettling finding from Shaw and Nave's research concerns confidence. Even when the AI was giving wrong answers roughly half the time, participants' confidence in their responses went up — access to AI inflated certainty across the board, regardless of whether that certainty was warranted.
At the heart of their paper is the concept of cognitive surrender: the tendency to adopt AI-generated answers with minimal scrutiny, allowing artificial cognition to override both intuition and deliberate reasoning. This is different from trusting a colleague's recommendation, because the AI's output carries no visible uncertainty, no hesitation, and no personal stake — all the normal social cues that ordinarily prompt us to push back.
For organizations, cognitive surrender is a workflow problem as much as an individual psychology problem. If the system is designed to make acceptance easy and disagreement slow, the outcome is structurally skewed toward acceptance before any deliberate review takes place. Recognizing this dynamic is the first step toward designing workflows that give genuine review a fighting chance.
What Counts as High Stakes?
Not every AI-assisted decision requires a deliberate pause. The practical question is how to classify decisions before designing oversight into them.
A high-stakes environment, in the context of AI use, refers to a setting in which AI systems operate under conditions where errors or failures could lead to serious, often irreversible outcomes — including harm to individuals, safety threats, substantial financial losses, legal liabilities, ethical breaches, or damage to public trust.
Irreversibility is the most reliable practical filter. A recommendation that can be undone in an afternoon belongs in a different category than a hiring decision, a pricing commitment, a contract term, or a clinical diagnosis. Senior executive guidance frames this explicitly: use reversibility and stakes as filters, and automate low-risk, reversible decisions first. The corollary is that high-consequence, difficult-to-reverse, and interpretive decisions need deliberate review points built into the workflow — not added as an afterthought.
A Practical Staging Model
Think of AI autonomy as a spectrum that should map to risk maturity rather than capability alone:
- Automated with logging — reversible, low-consequence, high-volume decisions where the AI's track record is well understood
- Human-in-the-loop — AI generates the recommendation; a human approves before any action is taken
- Human-on-the-loop — AI acts but a human monitors and can intervene; reserved for settings where predictability is high
- Human-led with AI support — AI surfaces information; humans own the reasoning and the decision entirely
Moving from human-in-the-loop to human-on-the-loop should happen only as predictability increases, not simply as AI capabilities expand or pressure to move faster grows. Mapping each decision category to a position on this spectrum before deployment — rather than after problems emerge — is a practical starting point for governance.
Automation Bias Does Not Respect Expertise
A reasonable assumption is that domain experts — experienced clinicians, seasoned lawyers, senior financial analysts — are less susceptible to automation bias than novices. The evidence does not fully support this.
As AI becomes increasingly embedded in high-stakes domains such as healthcare, law, and public administration, automation bias — the tendency to over-rely on automated recommendations — has emerged as a critical challenge in human–AI collaboration. The mechanism varies with expertise level in ways that make neither group immune.
Both familiarity and unfamiliarity with a task can lead to overreliance on AI, but through different pathways: experts may assume the AI handles routine decisions flawlessly and become complacent, while their confidence in their own understanding can cause them to overlook nuances the AI misses. Novices, meanwhile, lack the reference points needed to recognize when an AI output is implausible. Neither profile is safe by default.
This means that expertise cannot substitute for structural oversight. Organizations cannot rely on the judgment of experienced staff as their primary defense against automation bias if those staff are operating inside systems that penalize disagreement or obscure the evidence behind a recommendation. The implication is that safeguards must be designed into the system itself, not delegated to individual skill or vigilance.
Engineering the Pause: Cognitive Forcing Functions
If the problem is structural, the solution must be structural too. Cognitive Forcing Functions (CFFs) are interaction-level interventions used in human–AI collaboration to deliberately interrupt fast, heuristic reasoning and compel engagement with more deliberative, analytical processes at critical decision points. They aim to reduce overreliance on AI by requiring users to perform explicit reflective or analytic actions before accepting AI-generated suggestions.
One well-studied example is the staged answer reveal: the user must formulate an initial response before the AI suggestion is shown, and then reconcile the two. This small sequence change prevents the AI output from anchoring the human's thinking before any independent judgment has formed.
The research result is notable: cognitive forcing functions, which elicit analytical thinking, significantly reduced overreliance on AI compared to the simpler approach of presenting explanations for AI recommendations. Showing a user why the AI made a recommendation is less effective than requiring the user to commit to an independent judgment first.
Other CFF patterns organizations can encode into workflows include:
- Requiring explicit written justification before approving a high-stakes AI recommendation
- Surfacing a confidence score or uncertainty range alongside every AI output
- Introducing a mandatory cooling-off period before irreversible actions are confirmed
- Routing contested or outlier outputs to a second reviewer automatically
None of these interventions require replacing AI with human effort across the board. They are targeted friction — applied at the specific points where the cost of a wrong acceptance is highest.
The Ceremonial Oversight Problem
Designing oversight into a system and designing meaningful oversight are not the same thing. A person cannot preserve meaningful judgment if the workflow gives them thirty seconds to approve an AI recommendation, hides the underlying evidence, and penalizes them for slowing the process down. In that environment, human oversight becomes ceremonial — it exists on paper but has no practical influence on outcomes.
Enterprises need discipline encoded into workflows, permissions, evaluations, escalation rules, and system design. The problem is not only whether individual people remain thoughtful; the system must leave them enough evidence, authority, and time to think. Ceremonial oversight may satisfy a checkbox on a compliance form while providing no actual protection against consequential AI errors.
Practical tests for meaningful oversight:
- Does the reviewer have access to the underlying data and logic, not just the output?
- Is there a visible, low-friction path to reject or modify the recommendation?
- Is there a record of the reviewer's decision and rationale?
- Are reviewers evaluated partly on the quality of their challenges, not only on throughput?
If the answer to any of these is no, the oversight is likely ceremonial rather than substantive.
Governance Frameworks and Regulatory Pressure
Slow thinking is increasingly encoded into law, not just best practice.
In Europe, the requirements are broader in scope. Deployers of high-risk AI systems must implement human oversight measures as specified, ensure staff have AI literacy, conduct Fundamental Rights Impact Assessments before deploying in regulated sectors, maintain use logs for at least six months, and report serious incidents to the provider and national authorities. The EU AI Act's definition of high-risk covers a wide range of employment, credit, healthcare, education, and public administration use cases — which means these obligations apply to a large share of enterprise AI deployments.
Organizations that treat human oversight as a feature to add later, after deployment, are increasingly likely to find that regulatory frameworks treat it as a precondition. Building review structures before deployment is therefore both a governance best practice and, in many sectors, a legal requirement.
Rebuilding the Capacity for Independent Judgment
There is a longer-term concern underneath all of this. If workflows consistently allow AI to do the reasoning and humans to ratify the result, organizations may gradually erode the institutional capacity for independent judgment — the knowledge, heuristics, and critical instincts that made human reviewers valuable in the first place.
The reversibility benchmark applies not only to individual decisions but to the decision-making architecture itself: can the organization restore meaningful human decision-making if it chooses to? If the answer becomes no — because expertise has atrophied, documentation has thinned, and the rationale behind AI recommendations is opaque — then the organization has effectively made an irreversible decision by default, without ever consciously choosing to do so.
This is among the strongest arguments for building deliberate review into workflows from the beginning rather than retrofitting it after problems emerge. The cost of a meaningful pause before a consequential decision is almost always lower than the cost of discovering, after the fact, that the AI was wrong and no one was positioned to have caught it. Treating the organization's capacity for independent judgment as an asset that requires active maintenance is not a counsel against AI adoption — it is what makes AI adoption sustainable over time.
Putting It Together
The cognitive science, the governance literature, and the emerging regulatory frameworks converge on the same practical guidance:
- Classify decisions by reversibility and consequence before assigning AI autonomy levels
- Use cognitive forcing functions — especially staged answer reveals and explicit commitment requirements — to interrupt automatic acceptance at critical points
- Design oversight that is structurally meaningful: evidence visible, rejection path clear, rationale recorded
- Move toward greater AI autonomy only as predictability increases and the track record supports it
- Treat the organization's capacity for independent judgment as an asset that requires active maintenance
Speed is a genuine benefit of AI-assisted decision-making. But speed applied to the wrong decisions, or accepted without a structured pause, can convert a capability into a liability. The discipline of knowing when to slow down — and building that slowness into the system itself — is what separates AI that may improve outcomes from AI that merely accelerates them.
Where Oz fits
Oz by Anyreach is an AI teammate you can ask in Slack, Microsoft Teams, email, on the web, or on a call. Oz can use connected tools to help complete work, and every risky action asks for approval first. Learn how I Done This is transitioning to Oz.
Frequently asked questions
What is cognitive surrender and why does it matter for AI-assisted decisions?
Cognitive surrender, a term used by Wharton researchers Steven Shaw and Gideon Nave, refers to the tendency to adopt AI-generated answers with minimal scrutiny, allowing artificial cognition to override both intuition and deliberate reasoning. Research found that even when AI was wrong roughly half the time, participants' confidence in their responses rose — meaning users can become more certain while becoming less accurate.
How should organizations decide which decisions require human review?
Reversibility and consequence are the two most reliable filters. Low-risk, reversible decisions can be automated first. High-consequence, difficult-to-reverse decisions — such as those involving significant financial commitments, employment, healthcare, or legal exposure — should have deliberate review points built into the workflow before any action is taken.
Does domain expertise protect against automation bias?
Not reliably. Research indicates that both experts and novices are susceptible to automation bias, but through different pathways. Experts may assume the AI handles routine cases correctly and become complacent, while novices lack the reference points to recognize implausible outputs. Neither profile is safe without structural safeguards.
What are cognitive forcing functions and do they work?
Cognitive forcing functions are interaction-level interventions designed to interrupt fast, automatic thinking and require deliberate engagement before a user accepts an AI recommendation. One well-studied example is the staged answer reveal, where a user must form an independent judgment before seeing the AI's suggestion. Research found that cognitive forcing functions significantly reduced overreliance on AI compared to simply showing explanations for AI recommendations.
What is the difference between meaningful oversight and ceremonial oversight?
Ceremonial oversight exists on paper but has no practical influence — for example, when a reviewer has thirty seconds to approve a recommendation, cannot access the underlying evidence, and is penalized for slowing the process. Meaningful oversight requires that reviewers have access to the evidence behind a recommendation, a clear path to reject or modify it, a way to record their rationale, and enough time to engage genuinely.
What do current AI regulations require regarding human oversight?
The EU AI Act requires deployers of high-risk AI systems to implement specified human oversight measures, ensure staff AI literacy, conduct Fundamental Rights Impact Assessments before deploying in regulated sectors, maintain use logs for at least six months, and report serious incidents. Singapore's agentic AI framework independently requires organizations to define checkpoints that mandate human approval before high-stakes, irreversible, or atypical actions are executed.
Can an organization lose its capacity for independent judgment over time?
Governance literature suggests this is a genuine risk. If workflows consistently allow AI to do the reasoning while humans only ratify results, the expertise, critical instincts, and institutional knowledge that made human reviewers valuable may gradually atrophy. A practical test is whether the organization could restore meaningful human decision-making if it chose to — and whether that answer remains yes as AI becomes more deeply embedded.