The Big Problem
How would you feel if a doctor made an important decision about your care, and only afterward you learned that an AI system had already influenced the call? For many people, that wouldn’t feel like a small detail. We want to believe that high-stakes judgment still starts with a person—someone who can weigh nuance, read context, and respond to what doesn’t fit neatly on a screen. Yet across many settings, AI tools are increasingly being used to detect risk, rank urgency, and direct attention before a clinician or other frontline decision-maker has reached an independent view. That doesn’t mean these systems shouldn’t be used; in many cases, AI tools can improve speed, consistency, and efficiency. However, they may also make it harder to see where judgment begins and who owns it.
Healthcare is one example that brings that tension into focus, but it isn’t the only field facing it. Similar patterns may shape hiring, financial lending, insurance, the legal system, and public services, where AI can screen or flag people before a final decision is discussed. Experienced leaders already know what makes these tools attractive. They can help teams manage volume, reduce fatigue, and catch signals that could otherwise be missed. Even so, accountability doesn’t always hold once AI enters the process. A human may still approve the result, but it may no longer be clear who has the authority to question the system, override it, or answer for the outcome when the stakes are high.
TL;DR
- Organizations have integrated AI into consequential workflows faster than they’ve clarified accountability, leaving humans formally involved while authority, discretion, and responsibility become harder to locate.
- Delaying AI’s input and rewarding critical review helps preserve independent judgment, which makes it less likely that people will default to the system before forming an opinion of their own.
- Clearer lines of ownership, paired with better reporting and escalation methods, help organizations keep accountability from dissolving across a long and layered workflow.
What Is Human Accountability?
Human accountability is the clear assignment of responsibility for a decision and its consequences, even when AI helps inform the process. It means a person or clearly designated role still has the authority to review a recommendation, question it, override it, and decide whether to follow it. It also means responsibility can still be clearly identified if that decision is later challenged or causes harm.
When Human Judgment Meets Automated Decision Systems
Artificial intelligence has moved well past the pilot stage. For example, the Immigration, Refugees and Citizenship Department of Canada uses advanced analytics and automation to help process applications faster, and Canadian guidance already treats systems that generate assessments, recommendations, scores, or summaries for a human decision-maker as part of automated decision-making.1 In hiring, the legal pressure is building too. In July 2024, a federal judge allowed a discrimination case against Workday to proceed over allegations tied to AI-powered applicant screening.2 The case served as a warning that courts are paying closer attention to the role that software plays in shaping employment decisions.
Here’s where the accountability problem gets more concrete. A recruiter may still approve a shortlist, and a social worker might still sign off on a file, but software could already have ranked the applicants, flagged the client’s case, or directed attention toward one concern rather than another. In that kind of workflow, the organization can still point to the person who made the final call, yet it may be harder to see how much the system had been influencing that call before the review even began.
That’s why AI deployment isn’t only about choosing a tool. It’s also about deciding what the tool will do, how much weight people will give it, when they can depart from it, and who’s expected to answer if the result causes harm.
Newer rules are starting to reflect that problem more directly. Article 14 of the European Union’s AI Act says human oversight for high-risk systems should allow people to interpret outputs and decide whether to disregard, override, or reverse them.3 New York City’s Local Law 144 takes a similarly concrete approach in hiring: employers using automated employment decision tools must complete a recent bias audit, publish a summary of the results, and notify candidates.4
These rules also point toward a practical set of opportunities. Organizations can preserve human accountability by letting experienced decision-makers form an initial judgment before seeing the model’s recommendation. This rewards careful review rather than raw throughput, and adds small friction points such as brief justification prompts, override flags, or escalation triggers when stakes are high. Human accountability, then, cannot rest on a person’s presence alone. Instead, it rests on whether authority, discretion, and responsibility remain clear once AI has been built into the decision process.
Challenge #1: Human Oversight May Look Intact While Becoming Less Effective in Practice
A person may still be making the final call, but that doesn’t mean human judgment is working the way an organization expects. Once an AI system offers a diagnosis, a ranking, or a risk score, two familiar problems can start shaping the decision: automation bias and anchoring bias. Automation bias is the tendency to lean too heavily on automated advice, even when it’s flawed. Anchoring bias works a bit differently. The first recommendation sets the starting point, and everything that follows may be interpreted around it rather than from scratch.
In real workflows, these two biases can reinforce each other. The AI output arrives early, looks data-driven, and may begin framing the case before the person has fully worked through it alone. Recent reviews of AI-driven clinical decision support describe automation bias as a major implementation risk for exactly this reason.4,5
Healthcare offers one example of this pattern. In a study by Lyell and colleagues, 120 final-year medical students used an electronic prescribing system across simulated clinical scenarios.7 When the system’s guidance was inaccurate, participants became more prone to prescribing errors, and commission errors increased by 56.9% compared to situations without that incorrect support. A 2023 JAMA Network Open randomized vignette study found something similar at a larger scale: physician diagnostic accuracy improved when AI predictions were standard, but dropped by 11.3% when the AI predictions were systematically biased, and model explanations didn’t remove that effect.8
The problem, then, isn’t that humans vanish from the decision-making process. Rather, it’s that their role can gradually change from evaluator to confirmer once the system’s output becomes the starting point. That’s where anchoring bias becomes especially important; if the first recommendation is treated as the baseline, later reasoning may become an adjustment around the machine’s judgment rather than a fresh assessment of the case.
Outside healthcare, the same pattern shows up in public administration and more general decision-making. A 2024 legal analysis of automation bias highlights the Austrian Public Employment Service algorithm.9 The system was introduced to help caseworkers allocate resources by grouping jobseekers according to their predicted chances of labor market integration. Yet, the model penalized factors such as being a woman, being older, having health limitations, or coming from a non-European Union country. It was described as technical support or a second opinion, but Austrian regulators pointed to heavy caseloads and very limited time per case, conditions in which a meaningful second look couldn’t simply be assumed.
That’s a useful reminder that oversight can look solid on paper while being weak in practice. A 2024 experiment in Computers in Human Behavior found a related pattern in financially risky decisions: people overrelied on AI advice even when it conflicted with contextual information and worked against their own interests, and merely knowing that the advice came from AI increased reliance.10
What looks like human oversight may sometimes be little more than human confirmation. If AI enters the process early, is treated as the default, and shapes attention before a person has formed an independent view, responsibility may still be assigned on paper while real judgment has already narrowed.
behavior change 101
Start your behavior change journey at the right place
Opportunity #1: Preserve Human Accountability Through Timing and Incentives
If organizations want responsibility to remain visible after AI enters the picture, they may need to reconsider both when the tool appears and what people are encouraged to do once it does. When AI enters early, it may start framing what gets noticed before the reviewer has formed an initial view. Over time, that could leave oversight looking intact on paper while growing thinner in practice.
Let people form a view before AI frames the case
One practical response would be to change the order of the decision. Instead of letting the system speak first, organizations could require an initial human assessment before the model’s recommendation appears. A reviewer might record a short rationale, flag the factors they find most relevant, or make a provisional judgment before seeing the score or ranking.
In other settings, AI could be presented alongside other sources of information rather than appearing as the opening answer at the top of the screen. These are modest design choices, but they could help preserve the independence that accountability depends on. If people are being asked to evaluate a case, they need a fair chance to do that before the tool begins steering interpretation.
That instinct is already showing up in real institutions. Canada’s federal court has stated that it will not use artificial intelligence in making judgments and orders without first engaging in public consultation.11 It has also said that it will remain fully accountable to the public for any future use of AI in its decision-making work. This approach doesn’t ban AI forever, but it does show a disciplined sequencing. Core judgment is being held back until the surrounding conditions for responsibility have been considered. In accountability terms, that’s a strong example of refusing to let the tool become the first mover in a high-stakes decision before the human role has been properly defined.
Reward careful judgment instead of passive agreement
Even well-designed workflows may weaken if the organization rewards the wrong things. People usually learn very quickly what counts as good performance. If speed, throughput, and smooth compliance are treated as the main indicators of success, employees may start relying on AI in ways the organization never formally asked for. That’s why accountability has to be supported by incentives, not left as an abstract expectation.
Research on judgment has shown that people tend to reason more carefully when they know in advance that they’ll have to explain how they reached a decision.12 Studies of automation bias suggest something similar, indicating that reliance on automated advice increases under time pressure, while meaningful accountability for accuracy or performance can reduce automation-related errors.13 In practical terms, organizations could evaluate staff on the quality of their reasoning, require a brief justification for accepting or overriding an AI recommendation, and make escalation routes easy to use when something feels off.
Leaders could also reward well-grounded disagreement rather than treating it as delay or resistance. When people know that thoughtful challenge is part of the job, oversight is more likely to remain active, deliberate, and connected to a real human judgment rather than a rubber stamp.
Challenge #2: Responsibility Can Become Fragmented Across the Workflow
When an AI-assisted decision goes wrong, the obvious question is “Who’s responsible?” In many organizations, the answer isn’t easy to pin down. The model may have been built by one group, purchased by another, configured by someone else, and approved by a manager who only saw the final stage. Each person may have handled one piece of the process, yet no one fully owned the whole chain. That’s where responsibility can start to fragment, getting divided across the workflow in ways that make accountability harder to locate once harm occurs.
Diffusion of responsibility describes how people tend to feel less personally accountable when others are also involved. In AI settings, that tendency can become even stronger because layered workflows make it easier for each actor to assume that someone else has already checked the critical step. A frontline worker may think the tool has already been validated by technical experts, while a manager may believe the employee still owned the decision because they clicked “approve.” Each account may contain part of the truth, but the overall result is a process where ownership has become harder to trace.
The wrongful arrest of Robert Williams, driven by misidentification by facial recogntition technology, shows how that kind of fragmentation can harden into real, serious harm.14 The facial recognition system didn’t arrest Williams by itself—the damage emerged through a series of institutional steps. State police ran the facial recognition search while investigating a robbery. Detroit officers received the result as an investigative lead, and detectives used that lead to build a photo lineup. A witness identified Williams from that lineup, leading an officer to seek and obtain an arrest warrant.
By that point, the original error had already been passed forward and was treated as increasingly credible, as each step in the chain only validated the steps before. No single actor had to control the whole chain for the chain to produce harm anyway. The American Civil Liberties Union (ACLU) described the case as a wrongful arrest caused by reliance on a false face-recognition match, and the Detroit Police Department’s 2024 settlement imposed tighter limits, audits, and evidentiary requirements around the technology’s use. The remedy focused on the workflow because the workflow had been where responsibility broke apart.
A hiring system can produce the same problem even when software isn’t making the final choice on its own. In the previously-mentioned Workday litigation, the plaintiff alleged that the company’s tools were being used to score, sort, rank, or screen applicants for clients, while recruiters and hiring managers still handled later stages of review and selection.2 In July 2024, a federal judge allowed major claims to proceed on the basis that Workday, as an AI vendor, was subject to federal anti-discrimination laws because its tools were acting as an “agent” for its clients, performing part of a traditional hiring function.
This is an example of a decision path where ownership is difficult to localize. Workday can say it supplied the system. The employer can say it made the final hiring decision. Recruiters may say they reviewed only the pool that reached them. Yet an applicant could’ve been filtered out much earlier, before any downstream human had engaged with the file in a meaningful way. Human involvement is still present on paper, but responsibility has already been broken into segments that are easier to defend individually than to own collectively.
Then there’s the question of where blame lands after the damage is done. Madeleine Clare Elish, a prominent cultural anthropologist, offers an idea that’s especially helpful here: the moral crumple zone.15 In complex automated systems, the nearest human operator may end up absorbing the fallout even when they had limited control over how the larger system behaved. That can make organizations look accountable without making accountability accurate. For example, a frontline employee is easier to identify than a diffuse chain of designers, vendors, managers, procurement teams, and policy owners.
So, responsibility may not vanish at all. Instead, it could collapse onto the person who’s easiest to name after the fact. For senior leaders, that’s a governance problem as much as an ethical one. If responsibility has been scattered across procurement, design, implementation, review, and action, the organization may still be able to point to a human at the end of the process. What might be harder to show is that this person had enough authority, information, and control to deserve the full weight of accountability.
Opportunity #2: Connect What Different Parts of the Workflow Know
A stronger AI workflow depends on more than one person catching a problem. People across the process have to recognize failure modes, communicate what they’re seeing, and pass concerns along in ways others can act on.
Train people to spot real failure modes
Training should be more than teaching people which button to click. In an AI-enabled workflow, employees need to understand how the system tends to fail, which warning signs deserve attention, what their own role includes, and what the next person in the chain is likely to assume if a case gets passed forward. Case-based training is especially useful for that reason. Research on automation bias and complacency has found that exposing people to rare but meaningful automation failures during training can reduce complacency and improve checking behavior, even though it doesn’t remove overreliance entirely.16
People usually get better at spotting trouble when they’ve seen realistic bad cases rather than polished demos and ideal examples. A hospital using an AI diagnostic aid could walk radiologists and emergency clinicians through cases where the model performed poorly on unusual images, borderline presentations, or patient groups that had been underrepresented in development data. A hiring team could review cases where an automated ranking looked reasonable at first glance but had been driven by missing context, proxy variables, or patterns that screened out qualified candidates too early. A claims reviewer could be trained on examples where the recommendation aligned with the interface logic yet clashed with the medical record or policy context.
The goal isn’t to make staff fearful of the system, but to build a mental library of failure modes so that oversight becomes more informed and more specific. Effective review has always depended on domain judgment, and AI doesn’t erase that. In fact, it may be making that judgment more important.
Create small friction points that leave a usable trail
In many organizations, someone notices a recurring problem, handles it locally, and moves on, which may resolve the immediate case but often means the information stops there. A lightweight reporting method could help interrupt that pattern. If a user overrides an AI recommendation, they might select a short reason code such as “missing context,” “pattern mismatch,” or “confidence looks too low.” When the same code starts appearing repeatedly, it may be pointing to a deeper issue for the internal technical team, vendor, manager, or compliance lead.
If a downstream reviewer keeps encountering the same anomaly, that information shouldn’t remain buried in separate files or private workarounds. It should move upstream to the people who can examine the system, and across to teams that may be relying on the same tool under similar conditions.
Consider a housing benefits office, where caseworkers may keep overriding an AI-generated fraud risk score because the system flags applicants whose paperwork appears incomplete. Under the surface, the missing information reflects language barriers, unstable housing, or limited access to documents rather than likely misuse. Once each override is tagged with a reason code, the issue can move back to the team managing the model, reach supervisors who are tracking recurring concerns, and circulate across the office so other caseworkers aren’t left treating the same pattern as an isolated judgment call. What had looked like a one-off exception can start to register as a repeat problem, and that gives the organization a better chance of responding before the same error keeps resurfacing.
Small friction points can make that kind of communication much easier without turning the workflow into a slog. An organization doesn’t need a giant committee meeting every time someone questions a recommendation, though it may need a decision environment that makes it easier to pause, flag, and communicate when something looks off.
A short override prompt, a recurring anomaly flag, or an automatic notice after repeated reason codes may sound minor, yet they create a record others can actually use. That record connects downstream experience to upstream responsibility, and it gives technical teams, vendors, managers, and frontline users a better chance of seeing the same pattern before it keeps producing the same error. Then a reviewer’s concern becomes information the rest of the workflow can use, track, and respond to.
Caveats to Consider
Even the best accountability plan can run straight into ordinary organizational constraints. A leadership team may agree that AI oversight should be clearer, better documented, and easier to trace, yet the practical barriers are often less elegant than the policy language. Cost is one of those barriers; the Organisation for Economic Cooperation and Development (OECD) reported in its 2025 survey of more than 6,000 mid-level managers across France, Germany, Italy, Japan, Spain, and the United States that high cost was the leading reason firms gave for not adopting algorithmic management tools, cited by 79% of managers across countries.17
Managers in France, Germany, Italy, Japan, and Spain ranked staff resistance second, cited as a barrier by 56% to 67% of respondents. That doesn’t mean better oversight isn’t possible, but it does suggest that organizations may be trying to improve governance while staff are already worried about workload, trust, and whether the tool is worth the disruption.
Consultation can also become part of the same problem. In the same OECD survey, 94% of managers in the United States reported being consulted about algorithmic management tools, while only 32% said other employees were consulted more broadly.17 In Japan, those figures were 81% and 46%, respectively. If the people closest to day-to-day friction points feel they were brought in late or barely heard, resistance may harden and local workarounds may multiply. Then, the organization governs AI with an accountability method that looks coherent at the top and starts coming apart in daily use.
Preparing for an AI Future That Still Has Human Accountability
Accountability problems in AI rarely come from one dramatic failure. More often, they’ve been building inside ordinary workflows, where people encounter the tool too early, where responsibility has been divided across teams, and where one part of the process may spot a problem without the rest of the chain responding to it. Under those conditions, a human can still appear to be in control even though the practical basis for that control has already started to weaken. The real question is whether judgment, ownership, and communication have been built firmly enough to hold up once the system enters day-to-day use.
That pressure is likely to grow as more organizations try to introduce AI into routine operations, core services, and high-stakes decisions alike. Some of those choices carry financial, legal, operational, or human consequences, while others shape access, opportunity, safety, or trust more gradually over time. If accountability is going to remain credible, it can’t rest on policy language alone. It has to be built into the way work is carried out, through clear authority, usable reporting channels, and review processes people can follow under ordinary working conditions.
At The Decision Lab, we help organizations tackle the operational challenges that arise when innovation meets established practice. Whether the issue involves AI, a new policy, a redesigned process, or a broader organizational change, we apply decision science principles to understand how people actually make judgments, where friction builds, and which parts of a system are making good outcomes harder to achieve. If your team is considering how AI should be introduced across real working environments, we’d be happy to help shape an approach that supports stronger performance while preserving trust, judgment, and clear lines of ownership.
Related TDL articles
Why Do We Keep Handing More Decisions to AI?
Using AI to brainstorm one difficult email, then gradually letting it draft even routine replies you could easily write yourself, is a pretty ordinary example of how delegation creep starts. Delegation creep is the tendency for people and organizations to steadily expand what they let automated systems do, often moving from small conveniences into judgments with real strategic, ethical, or human consequences. This piece looks at the psychological, behavioral, and social forces that make that pattern so easy to miss, and what can be done before it goes too far.
The Dangers of an Artificially Intelligent Future
AI can save time and extend human capacity, but it can also learn from biased datasets shaped by racism, sexism, and other entrenched inequalities, then reproduce those patterns in outputs that seem neutral. So how intelligent is our artificially “intelligent” future, really, and how do we make it more careful, more accountable, and less likely to carry old prejudices forward? This article examines where those risks come from, why they’re so often underestimated, and how AI can be used with stronger safeguards.
Sources
- Immigration, Refugees and Citizenship Canada. (2026, February 20). How we use advanced analytics, automation and other technologies. Government of Canada. https://www.canada.ca/en/immigration-refugees-citizenship/corporate/transparency/digital-transparency-advanced-data-analytics.html
- Wiessner, D. (2024, July 16). Workday must face novel bias lawsuit over AI screening software. Reuters. https://www.reuters.com/legal/litigation/workday-must-face-novel-bias-lawsuit-over-ai-screening-software-2024-07-15/
- EU Artificial Intelligence Act. (2026, August 2). Article 14: Human oversight. https://artificialintelligenceact.eu/article/14/
- New York City Department of Consumer and Worker Protection. (n.d.). Automated employment decision tools (AEDT). NYC.gov.
- Abdelwanis, M., Alarafati, H. K., Tammam, M. M. S., & Simsekler, M. C. E. (2024). Exploring the risks of automation bias in healthcare artificial intelligence applications: A Bowtie analysis. Journal of Safety Science and Resilience, 5(4), 460–469. https://doi.org/10.1016/j.jnlssr.2024.06.001
- Khera, R., Simon, M. A., & Ross, J. S. (2023). Automation bias and assistive AI: Risk of harm from AI-driven clinical decision support. JAMA, 330(23), 2255–2257. https://doi.org/10.1001/jama.2023.22557
- Lyell, D., Magrabi, F., & Coiera, E. (2019). Reduced verification of medication alerts increases prescribing errors. Applied Clinical Informatics, 10(1), 66–76. https://doi.org/10.1055/s-0038-1677009
- Jabbour, S., Fouhey, D., Shepard, S., Valley, T. S., Kazerooni, E. A., Banovic, N., Wiens, J., & Sjoding, M. W. (2023). Measuring the impact of AI in the diagnosis of hospitalized patients: A randomized clinical vignette survey study. JAMA, 330(23), 2275–2284. https://doi.org/10.1001/jama.2023.22295
- Allhutter, D., Cech, F., Fischer, F., Grill, G., & Mager, A. (2020). Algorithmic profiling of job seekers in Austria: How austerity politics are made effective. Frontiers in Big Data, 3, Article 5. https://doi.org/10.3389/fdata.2020.00005
- Klingbeil, A., Grützner, C., & Schreck, P. (2024). Trust and reliance on AI: An experimental study on the extent and costs of overreliance on AI. Computers in Human Behavior, 160, 108352. https://doi.org/10.1016/j.chb.2024.108352
- Federal Court. (2024, May 7). Notice to the parties and the profession: The use of artificial intelligence in court proceedings. https://www.fct-cf.ca/Content/assets/pdf/base/FC-Updated-AI-Notice-EN.pdf
- Lerner, J. S., & Tetlock, P. E. (1999). Accounting for the effects of accountability. Psychological Bulletin, 125(2), 255–275. https://doi.org/10.1037/0033-2909.125.2.255
- Goddard, K., Roudsari, A., & Wyatt, J. C. (2012). Automation bias: A systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1), 121–127. https://doi.org/10.1136/amiajnl-2011-000089
- American Civil Liberties Union. (2024, January 29). Williams v. City of Detroit. https://www.aclu.org/cases/williams-v-city-of-detroit-face-recognition-false-arrest
- Elish, M. C. (2019). Moral crumple zones: Cautionary tales in human-robot interaction. Engaging Science, Technology, and Society, 5, 40–60. https://doi.org/10.17351/ests2019.260
- Bahner, J. E., Hüper, A.-D., & Manzey, D. (2008). Misuse of automated decision aids: Complacency, automation bias and the impact of training experience. International Journal of Human-Computer Studies, 66(9), 688–699. https://doi.org/10.1016/j.ijhcs.2008.06.001
- Milanez, A., Lemmens, A., & Ruggiu, C. (2025, February). Algorithmic management in the workplace: New evidence from an OECD employer survey (OECD Artificial Intelligence Papers No. 31). OECD. https://www.oecd.org/content/dam/oecd/en/publications/reports/2025/02/algorithmic-management-in-the-workplace_3c84ed6d/287c13c4-en.pdf
















