Why do we keep handing more decisions to AI?

Delegation creep in AI is a bias where people and organizations steadily expand what they let automated systems do.

Where this bias occurs

It often starts with low-stakes tasks like sorting emails or drafting boilerplate, then drifts into screening candidates, triaging welfare cases, or shaping messages with real strategic or ethical weight. Human factors research shows related patterns of automation bias and complacency, where decision makers miss problems the system does not flag and follow incorrect recommendations after good past performance. Over-reliance grows with increasing time pressure and distraction, so sensitive decisions end up flowing through systems never designed or governed for that level of responsibility.

Delegation creep appears wherever AI tools are embedded into workflows as helpers or copilots. In personal productivity, people start by using an assistant to manage calendar invites and inbox triage, then drift toward letting it prioritize which opportunities or relationships receive attention first. In content creation, teams adopt AI to write outlines and rough drafts, then lean on it to choose angles, frame arguments, and even decide which audiences to prioritize.

In operational settings, the pattern is sharper. Research on automation in complex systems shows that once a decision aid is introduced, operators quickly shift monitoring and judgment toward the automated channel.1 Under routine conditions, they tend to assume the system is functioning correctly and may overlook cues that contradict its output.2 Studies of trust in automated systems find that when tools are framed as accurate and efficient, people become more willing to lean on them and less inclined to question their recommendations, especially under time pressure.3

Policy and ethics work on AI governance has flagged a similar problem with formal oversight requirements. Many regulations call for human review of automated decisions, yet in practice, “human-in-the-loop” often means adding a signature step at the end of a process that is otherwise automated.4 When those humans lack time, context, or real authority to override the system, review becomes a ritual rather than a safeguard. The concept of moral crumple zones captures how humans can be left to absorb blame when things go wrong, even though the structure of the system gave them little real control in the first place.5

In workplaces governed by algorithmic management, automated systems assign tasks, rate performance, and suggest disciplinary steps. Empirical Human-Computer Interaction (HCI) research shows that such tools often begin as scheduling aids or performance dashboards and gradually become central to decisions about pay, promotion, and continued employment.6 Employer surveys indicate that once in place, these systems are frequently expanded to new uses without a corresponding upgrade in governance or worker voice.7 Organizational psychology research links this dynamic to reduced autonomy and motivation, as employees adapt their behavior to what the system measures and rewards rather than to professional judgment.8 Labor law scholars warn that, without safeguards, this expansion of delegated decision making can entrench discrimination and weaken existing protections at work.9

Individual effects

For an individual professional, delegation creep usually begins as relief. The AI assistant takes care of things that used to consume the workday. A drafting tool writes first passes for routine documents. A recommender ranks cases by urgency.

At first, people check the outputs carefully. They compare the AI’s triage to their own sense of priority. They edit drafts line by line. They scan recommended cases to spot anything that looks off. Over time, two related forces pull in the same direction. Repeated use creates habits that make it feel natural to let the system handle more work. Repeated positive outcomes create trust that the system is accurate and safe. Experiments on automation bias show that people tend to rely on automated aids, especially when they have performed well in earlier trials.1

As trust grows, the mental model shifts. The AI becomes “how things are done.” Warnings that “AI can make mistakes” turn into background noise when most outputs look reasonable. Monitoring becomes more selective. People focus on edge cases or rare alerts and skim past ordinary outputs. Automation complacency studies describe this as reduced vigilance in monitoring automated systems, particularly in cases of multitasking and heavy workload.2

In that environment, it feels natural to ask more of the tool. A manager who has used AI to rephrase feedback might begin asking it to suggest performance ratings. A clinician who uses decision support for dosage checks might lean on it for diagnostic ranking. A team lead who uses an assistant to summarize candidate materials might start relying on it for shortlist recommendations.

Delegation creep also interacts with trust and risk perception. Survey-based studies on automation bias and aversion suggest that when people perceive an automated system as accurate and efficient, they are more willing to lean on it, even in situations where they would normally hesitate.3 Once a tool has built a track record of apparently good performance, many users feel that questioning it is second-guessing an expert.

Systemic effects

Delegation creep changes organizations as much as it changes individuals. In algorithmic management environments, automated dashboards and scoring systems frame what managers see and what they treat as important. A conceptual explainer from Data & Society describes algorithmic management as a set of tools that structure working conditions and mediate how decisions about workers are made.6 When those tools begin by allocating shifts and transition to recommending terminations or bonuses, delegation creep has already occurred.

An OECD survey of employers found that AI systems used for scheduling, monitoring, and evaluation often expand in function over time, sometimes without clear new governance or worker input.7 A workforce might experience this as a steady shift from human supervisors who know their context toward opaque scorecards that are difficult to contest. Managers, for their part, may feel that they are simply “following the system” rather than making independent calls.

Organizational psychology research on algorithmic management suggests that heavy reliance on these tools can undermine basic needs for autonomy and competence at work, especially when metrics and recommendations are treated as unquestionable.8 Delegation creep then affects culture. People learn that pushing back against AI outputs is risky or futile. They adjust their behavior to fit what the system rewards.

Ethics work on AI governance adds another layer. The concept of moral crumple zones describes situations where humans absorb blame when something goes wrong, even though a complex automated system constrained their decisions.5 Delegation creep creates fertile ground for such zones. Over time, humans become the face of decisions that are functionally driven by AI recommendations and rules. When harm occurs, their signatures are easy to point to, while the slow shift of judgment to the system remains invisible.

Why it happens

Delegation creep does not require careless people or reckless organizations. It grows from structural incentives, cognitive habits, and the way AI tools are framed.

Convenience and cognitive load

AI systems promise efficiency. In real workflows, they deliver it. Once a tool reliably saves time on low-stakes tasks, it is practical to rely on it more frequently. Under time pressure, it becomes reasonable to check on it less. Studies of automation complacency show that operators monitoring automated systems perform worse on manual tasks when they try to monitor everything at once.2 Delegating more to the system can feel like the only way to cope.

Trust built from early success

Early interactions that go well raise perceived reliability. Experimental work on automation bias and algorithm reliance finds that people use past performance as a strong cue when deciding how much to trust a system.1 When the AI has handled a hundred routine cases, seemingly without issue, skepticism feels less justified. The leap from checking everything to checking spot cases feels small in the moment.

Framing of AI as “assistant” or “copilot”

The language around AI tools matters. When a system is described as an assistant that “takes care of the busywork,” it frames technical limits as a background concern rather than a central issue. Research on algorithmic management notes that employers often present automated systems as neutral, data-driven helpers that rationalize decisions.6 That narrative lowers perceived risk in extending their reach.

Shallow human oversight requirements

Legal and policy frameworks increasingly require human oversight of automated decision-making. A critical analysis of “human in the loop” policies argues that they often specify human presence without ensuring meaningful control.4 If oversight consists of clicking “approve” on AI outputs for dozens of cases per hour, the structure pushes toward rubber-stamping. Delegation creep thrives in the gap between the appearance of control and its reality.

Diffused responsibility in complex organizations

In large institutions, different teams own different parts of AI systems. Technical staff build models. Product teams design workflows. Compliance signs off on documentation. Managers implement tools on the ground. Work on responsibility in human-machine systems shows how this division can create moral crumple zones where no one actor feels wholly accountable.5 Delegation creep leverages that structure. As AI takes on more judgment, everyone can point to someone else for ultimate responsibility.

behavior change 101

Start your behavior change journey at the right place

Why it matters

Delegation creep changes which tasks AI performs without a clear debate about whether those tasks are appropriate for automation. At the technical level, models often perform very differently across task types and contexts. A tool that handles routine classification well may struggle with ambiguous edge cases or situations that require contextual judgment. When delegation creep moves more of those cases into the automated channel, errors accumulate in the very places where human judgment is most needed.

At the ethical level, expanding AI roles without explicit consent can undermine trust. Workers who agreed to algorithmic scheduling may find that the same system now informs performance evaluations and disciplinary decisions. Welfare recipients who expected automated checks for obvious errors may discover that substantive judgments about eligibility now flow through opaque models.

At the political level, delegation creep can shift where power sits in a system. If teachers, caseworkers, or clinicians feel that they mainly execute what AI systems recommend, their professional autonomy shrinks. That change in power is often invisible in official charts. On paper, people still make the decisions. In practice, their scope to disagree narrows.

AI policy debates increasingly stress the need for meaningful human control, especially in high-impact domains. Analyses of human oversight and AI regulation caution that nominal human involvement cannot substitute for real authority and capacity to intervene.4 Delegation creep is one of the main reasons that distinction matters. It is possible to have many names on a process and minimal real human judgment at key steps.

How it affects product

For product teams and service organizations, delegation creep creates a tension between scale and safety. Delegating more decisions to AI can improve throughput and consistency. A customer support team that uses AI to draft replies and classify tickets can serve more people quickly. A benefits office that uses scoring models to triage cases can allocate scarce staff time to what looks most urgent. An HR department that uses AI to pre-screen applications can reduce manual workload on recruiters.

The commercial logic is clear. Algorithmic management explains how organizations use automated decision systems to rationalize and streamline work, often under pressure to do more with less.6 Yet as tools move from classification and drafting toward recommendation and decision, the risk profile changes. Errors are no longer confined to low-stakes outputs. They affect who gets hired, which complaints are escalated, and which claims receive deeper scrutiny.

An OECD employer survey found that organizations often add new AI use cases on top of existing ones without fully revisiting governance structures.7 Teams that once used analytics to monitor trends begin relying on similar models for individual-level decisions. Each incremental step seems reasonable. Together, they create a system in which AI plays a central, sometimes opaque, role in shaping outcomes.

Worker-centered research on algorithmic management points out that over-reliance on automated systems can reduce creativity, adaptability, and motivation, especially when people feel monitored by tools they do not control.8 Delegation creep contributes by turning AI from a tool that supports human judgment into an invisible manager whose outputs are rarely questioned.

From a product perspective, the danger is reputational as well as operational. A system that was marketed as an assistant can become a source of scandal when it is revealed that critical decisions were effectively outsourced to an AI model with little oversight.

How to avoid it

Mitigating delegation creep requires purposeful boundaries, real oversight, and feedback loops that check where AI is actually in use.

Define task tiers explicitly

Organizations can map tasks into tiers: low-stakes routine, medium-stakes judgment, and high-stakes critical decisions. For each tier, they can specify which tasks AI may perform alone, which require human review with clear criteria, and which must remain human-led. This turns “we use AI to help” into concrete rules.

Set hard red lines

In some domains, certain decisions should not be delegated to AI at all, regardless of performance. That might include final hiring decisions, benefit cancellations, or high-risk clinical interventions. Explicit red lines help prevent scope creep that feels natural in the moment but crosses ethical or legal boundaries.

Design oversight for depth

Human-in-the-loop structures need to give reviewers time, information, and authority to disagree with AI outputs. Analysts of oversight practices argue that human review must be meaningful.4 That can involve sampling strategies, escalation routes, and metrics that reward catching AI errors rather than punishing slower throughput.

Monitor drift in practice

Delegation creep often happens informally. Teams discover novel uses for tools and adopt them without formal approval. Periodic audits and qualitative interviews can surface where AI is already shaping decisions. Findings can then feed back into governance and training.

Train for healthy skepticism

Training should frame AI systems as powerful but limited tools. Human factors research suggests that awareness of automation bias and complacency can help operators maintain more balanced trust in technology.1 This does not eliminate bias, but it provides a language and expectation for questioning outputs, especially when the stakes are high.

Give affected people channels to contest decisions

Delegation creep has real consequences for those on the receiving end of AI-influenced decisions. Clear appeal processes, explanation mechanisms, and support for contesting outcomes can help surface harms and boundary breaches early. They also affirm that responsibility ultimately rests with institutions.

Example 1 – Welfare automation and the Robodebt scheme

In Australia, the Robodebt scheme illustrates how delegation creep can unfold in public administration. The program, introduced in the mid 2010s, used automated data matching between tax records and welfare payment information to identify possible overpayments. When the system detected a discrepancy, it often generated debt notices to recipients based on averaged income data.

Over time, the automated process became central to compliance operations. Staff who had previously investigated discrepancies manually were expected to rely on system outputs. Recipients were told that they needed to disprove debts calculated by the algorithm. Many reported confusion and distress when confronted with large debts that they did not understand or believe they owed.

A Royal Commission and subsequent analyses concluded that the scheme was unlawful and that it represented a major failure of public administration.10 Automation had moved from supporting staff to effectively determining who was targeted for compliance action and how debts were calculated. Human oversight existed in theory. In practice, staff and managers relied heavily on automated outputs, and structural pressures discouraged them from challenging the system.

Delegation creep led a tool designed to streamline calculations to take on the effective power to accuse people of fraud and trigger aggressive debt collection. When the harms became undeniable, officials and agencies struggled to explain who had made the decision to rely so heavily on automation for such high-stakes judgments.

Example 2 –  Legal research and the Mata v. Avianca case

In 2023, a case in the U.S. District Court for the Southern District of New York drew global attention when lawyers were sanctioned for submitting a brief that relied on fabricated case citations generated by ChatGPT.11

According to court records and reporting, the attorneys used ChatGPT to help draft legal arguments and generate supporting case law. The system produced plausible-sounding citations and quotations that referred to non-existent decisions. When opposing counsel and the judge could not locate several cases, the court ordered the lawyers to provide copies. They again turned to ChatGPT, which confidently asserted that the cases were real and available in legal databases.

The lawyers later told the court that they had been unaware that ChatGPT could invent cases. They had not verified the citations through standard legal research tools before filing. The judge imposed a monetary sanction and emphasized that relying on an AI system did not absolve attorneys of their professional duty to ensure the accuracy of submissions.

This episode is a vivid example of delegation creep at the individual level. AI may have entered the workflow as a drafting aid, a way to save time on language and structure. It ended up taking over core professional responsibilities, such as identifying relevant precedent and validating citations. Oversight shifted from careful checking to trusting that a tool that wrote coherent text would also produce reliable legal references.

The case raised broader questions for the legal profession. How far should lawyers go in delegating research and reasoning to AI? What counts as acceptable use versus abdication of judgment? Mata v. Avianca shows how quickly the scope of delegation can slide when a system feels competent, and the pressures of workload and efficiency are real.

Summary

What it is

Delegation creep in AI occurs when people and organizations gradually expand the range of tasks they hand to AI systems, moving from low-stakes support work toward complex, judgment-heavy, or high-impact decisions. The shift often happens without explicit discussion or new safeguards, so oversight thins out while reliance grows.

Why it happens

Early success, time pressure, and positive framing make delegation feel natural. Psychological patterns such as automation bias and automation complacency encourage people to trust tools that have performed well in the past. Structural features such as shallow “human in the loop” requirements and diffused responsibility make it easy for AI to quietly take on more judgment while everyone assumes that someone else is watching closely.

Example #1 - Welfare automation and Robodebt

Australia’s Robodebt scheme used automated data matching and income averaging to generate welfare debt notices at scale. Over time, staff and managers relied heavily on system outputs, and recipients were expected to dispute automated debts. A Royal Commission later found the scheme unlawful and described it as a major public administration failure. The episode shows how welfare automation drifted from support to a de facto decision maker, leaving human oversight paper-thin.

Example #2 - Legal research in Mata v. Avianca

In Mata v. Avianca, lawyers used ChatGPT to generate case citations and quotations, then filed a brief containing entirely fictitious precedents. The court sanctioned them after they failed to verify the AI’s outputs through standard research tools. This case illustrates how a generative AI tool that begins as a drafting aid can end up taking over core professional judgment when users assume that fluent language implies reliable content.

How to avoid it

Organizations can limit harmful delegation creep by defining clear task tiers, setting hard red lines for decisions that must remain human-led, and designing oversight that gives reviewers time and authority to challenge AI outputs. Regular audits of how tools are actually used can catch informal scope expansions. Training that builds healthy skepticism and appeal channels that let affected people contest AI-influenced decisions both help keep responsibility anchored in people and institutions that can explain and correct what these systems do.

Related TDL articles

Automation bias

Why do people lean so heavily on automated advice, even when it is imperfect? This article examines automation bias in fields like aviation, healthcare, and finance, and explores design choices that help humans stay engaged rather than rubber-stamping what the system says.

Self-serving bias

How does our tendency to claim credit for success and deflect blame for failure play out in data-driven organizations? This article looks at self-serving bias in AI projects and offers ways to build cultures that encourage owning mistakes.

Sources

  1. Skitka, L. J., Mosier, K., & Burdick, M. (1999). Does automation bias decision-making. International Journal of Human-Computer Studies, 51(5), 991–1006. https://doi.org/10.1006/ijhc.1999.0252
  2. Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410. https://doi.org/10.1177/0018720810376055
  3. Gsenger, R., Eier, J., Csillag, J., & Schlogl, S. (2021). Trust, automation bias and aversion: Investigating trust in automated systems. Interdisciplinary Description of Complex Systems, 19(4), 542–560.
  4. Green, B. (2022). The flaws of policies requiring human oversight of artificial intelligence. Computer Law & Security Review, 46, 105710. https://doi.org/10.1016/j.clsr.2022.105681
  5. Elish, M. C. (2019). Moral crumple zones: Cautionary tales in human–robot interaction. Engaging Science, Technology, and Society, 5, 40–60. https://doi.org/10.17351/ests2019.260
  6. Lee, M. K., Kusbit, D., Metsky, E., & Dabbish, L. (2015). Working with machines: The impact of algorithmic and data-driven management on human workers. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (pp. 1603–1612). https://doi.org/10.1145/2702123.2702548
  7. Milanez, A., Lemmens, A., & Ruggiu, C. (2025). Algorithmic management in the workplace: New evidence from an OECD employer survey (OECD Artificial Intelligence Papers, No. 31). OECD Publishing. https://doi.org/10.1787/287c13c4-en
  8. Gagné, M., Parent-Rocheleau, X., Bujold, A., Gaudet, M.-C., & Lirio, P. (2022). How algorithmic management influences worker motivation: A self-determination theory perspective. Canadian Psychology / Psychologie canadienne, 63(2), 247–260. https://doi.org/10.1037/cap0000324
  9. De Stefano, V., & Taes, S. (2023). Regulating AI at work: Labour relations, automation, and discrimination. International Journal of Comparative Labour Law and Industrial Relations, 39(1), 13–40. 
  10. Shah, C. (2023, July 26). Australia’s Robodebt scheme: A tragic case of public policy failure. Blavatnik School of Government, University of Oxford. https://www.bsg.ox.ac.uk/blog/australias-robodebt-scheme-tragic-case-public-policy-failure
  11. Merken, S. (2023, June 26). New York lawyers sanctioned for using fake ChatGPT cases in legal brief. Reuters.https://www.reuters.com/legal/new-york-lawyers-sanctioned-using-fake-chatgpt-cases-legal-brief-2023-06-22/

About the Author

White guy wearing a white lab coat over a baby blue dress shirt.

Adam Boros

Researcher, Mount Sinai Hospital

Adam studied at the University of Toronto, Faculty of Medicine for his MSc and PhD in Developmental Physiology, complemented by an Honours BSc specializing in Biomedical Research from Queen's University. His extensive clinical and research background in women’s health at Mount Sinai Hospital includes significant contributions to initiatives to improve patient comfort, mental health outcomes, and cognitive care. His work has focused on understanding physiological responses and developing practical, patient-centered approaches to enhance well-being. When Adam isn’t working, you can find him playing jazz piano or cooking something adventurous in the kitchen.

About us

We are the leading applied research & innovation consultancy

Our insights are leveraged by the most ambitious organizations

Image

“

I was blown away with their application and translation of behavioral science into practice. They took a very complex ecosystem and created a series of interventions using an innovative mix of the latest research and creative client co-creation. I was so impressed at the final product they created, which was hugely comprehensive despite the large scope of the client being of the world's most far-reaching and best known consumer brands. I'm excited to see what we can create together in the future.

Heather McKee

BEHAVIORAL SCIENTIST

GLOBAL COFFEEHOUSE CHAIN PROJECT

OUR CLIENT SUCCESS

$0M

Annual Revenue Increase

By launching a behavioral science practice at the core of the organization, we helped one of the largest insurers in North America realize $30M increase in annual revenue.

0%

Increase in Monthly Users

By redesigning North America's first national digital platform for mental health, we achieved a 52% lift in monthly users and an 83% improvement on clinical assessment.

0%

Reduction In Design Time

By designing a new process and getting buy-in from the C-Suite team, we helped one of the largest smartphone manufacturers in the world reduce software design time by 75%.

0%

Reduction in Client Drop-Off

By implementing targeted nudges based on proactive interventions, we reduced drop-off rates for 450,000 clients belonging to USA's oldest debt consolidation organizations by 46%

Notes illustration

Eager to learn about how behavioral science can help your organization?