Obedience

What is Obedience?

Obedience is the psychological process of following explicit instructions or commands from an authority figure, even when those orders conflict with our values, morals, or sense of right and wrong. In everyday life, obedience ensures societal cohesion: following traffic lights, respecting workplace protocols, or adhering to school rules all hinge on it. Obedience is a powerful force that shapes behavior, sometimes with profound consequences.

The Basic Idea

One late afternoon in a bustling downtown hospital, the air buzzed with urgency. Monitors beeped, stretchers wheeled past, and nurses huddled over charts. Amid the chaos, a junior nurse paused by the medication station. She’d just noticed a senior physician, white coat crisp, stethoscope slung with quiet authority, preparing an unusual drug combination for a cardiac patient. Something about the dosage felt off. Her training flagged a warning, but so did her instincts. Still, she hesitated. The doctor was a specialist with decades more experience. Surely he knew better? She opened her mouth. Then closed it. “He must know what he’s doing,” she thought, turning back to her tasks. Minutes later, the patient coded.

That split-second decision wasn't unique; it reflects a powerful, often overlooked force in human behavior: obedience. Under the weight of perceived authority, our internal moral compass can be switched off, even in life-and-death situations. At the core of obedience lies a psychological transformation known as the agentic shift: individuals see themselves not as independent moral agents but as instruments carrying out orders. This shift emerged in Stanley Milgram's landmark experiments. Ordinary people, under the influence of perceived authority, administered what they believed were increasingly harmful electric shocks, some reaching the maximum voltage, despite inner turmoil and discomfort.1 This obedience wasn’t compliance with peers or passive conformity to social norms; it was active submission to commands from someone in power.2

Milgram’s findings show that obedience is neither blind nor weak-mindedness, but a consequence of context and perceived legitimacy. His participants, ranging in age and background, repeatedly overrode personal ethics when prodded by a stern experimenter.¹ Follow-up studies confirmed that weakening authority cues, by moving the setting from a prestigious university to a dingy office, dropped obedience rates, demonstrating how context amplifies compliance.3

This tendency isn’t just a historical artifact; it has real-world implications. During trials after World War II, SS officer and holocaust organizer Adolf Eichmann defended his actions by insisting he was just following orders—a chilling echo of Milgram’s conclusions.4 In hospitals, junior staff may remain silent when noticing errors due to hierarchical pressure. In corporate environments, employees may pursue unethical strategies because a superior sanctioned them. When obedience becomes automatic, accountability vanishes, and harmful decisions pass through unnoticed until consequences emerge.

Obedience doesn’t operate alone; it’s triggered by specific authority cues, white coats, titles, uniforms, or formal rituals that signal legitimacy. Once a directive is issued, we enter an agentic state, transferring responsibility from ourselves to the authority figure. The obedience cycle concludes when we take actions we wouldn’t choose independently, helping to explain how ordinary people follow orders they might personally reject.

Understanding obedience this way is more than academic—it’s a roadmap for designing interventions. By identifying authority cues and interrupting agentic shifts, institutions can foster environments where people balance respect for authority with personal responsibility. The challenge lies not in dismantling authority, but in teaching individuals when to question the command and stand by their moral instincts.

“Ordinary people, simply doing their jobs, and without any particular hostility on their part, can become agents in a terrible destructive process. […] The most common characteristic of those obeying is not sadism or aggression, but a lack of personal responsibility.”


— Stanley Milgram, Social Psychologist1

Key Terms

Agentic State: The state where individuals stop seeing themselves as personally accountable and instead act as instruments of someone else’s will, justifying their actions as “just following orders”. In this state, they feel less guilt or moral conflict because responsibility has been outsourced. It’s a crucial mechanism for understanding how otherwise moral people can cause harm when acting under direction.

Agentic Shift: The agentic shift is the psychological process where a person stops acting on their own judgment and begins to follow the directives of an authority figure. It’s the transition into the agentic state, where responsibility feels transferred to someone higher up. This shift reframes the self not as a decision-maker, but as a tool carrying out orders.

Perceived Legitimacy: Obedience thrives when the authority figure appears valid. This perception is shaped by context, symbols (like lab coats or titles), and environment. Whether it’s Milgram’s prestigious Yale lab or a hospital setting, the more legitimate the figure seems, the more likely people are to comply, even against their better judgment.

Authority Cues: The subtle, often non-verbal signals indicating that someone holds power, such as uniforms, seating positions, vocal tone, or technology interfaces. They act as behavioral triggers, priming individuals to follow instructions without critical thought. Across domains like healthcare and AI, authority cues help explain how obedience is initiated without explicit force.

Obedience Engineering: A modern term referring to how institutions or technologies design structures that intentionally influence compliance. From classroom behavior games to AI friction tools, this concept frames obedience not as accidental, but as something that can be constructed, manipulated, or resisted by design. It's especially relevant in education, software, and policy.

History

The Second World War and the Holocaust brought obedience into sharp moral glare. Villains like Adolf Eichmann claimed they were mere employees, “carrying out orders.” In 1961, social psychologist Stanley Milgram at Yale read the transcripts of Eichmann’s trial and wondered: could “ordinary people”, those who might teach a child, drive a neighbor, share a laugh, end up harming another human on command? He recruited volunteers for what was described as a memory experiment. Unbeknownst to them, the shocking twist played out behind a curtain. The participants were instructed to deliver electric shocks to a person in another room—someone they believed was a real volunteer—each time that person gave a wrong answer in a word-pair memory test. A researcher dressed in a lab coat stood beside them, calmly urging them to raise the shock level after each mistake. Through the wall, they heard the learner scream in pain, plead to stop, and eventually fall silent. Still, more than half continued to the highest voltage labeled on the machine, a level marked as dangerous and—if real—potentially fatal.1

Milgram’s original results shocked the world. But he didn’t stop there. He began tweaking the setup, running dozens of variations. In one version, everything stayed the same—same script, same commands—but the location changed. Instead of a sleek Yale lab, the experiment took place in a rundown office with scuffed walls and flickering lights. Obedience dropped.2 In another version, the researcher gave instructions over the phone instead of in person. Without the lab coat and direct eye contact, many participants disobeyed or pretended to obey while secretly skipping shocks.5 When the learner’s voice could be heard through the wall, more people hesitated. When the learner sat just a few feet away, fewer participants went to the end. And when the participant had to physically place the learner’s hand on a metal plate to deliver the shock, most refused. Across all these versions, one pattern emerged: obedience didn’t depend on personality, but on distance and symbolism. The closer the victim, the lower the obedience. The more distant the authority, the weaker their grip.

While Milgram’s lab conjured theoretical fears, the 1966 hospital experiment by psychiatrist Charles K. Hofling asked real healthcare professionals to test obedience in context. Through a phone call from a non-existent Dr. Smith, nurses were ordered to administer a double-dose of a fictitious drug named Astroten. Despite rules forbidding phone orders and overdosing, 21 of 22 nurses complied.5 Protocols and training were swept aside by a disembodied voice of authority. The conclusion wasn’t academic—it was a moral shock: obedience persists even among professionals trained to care.

In 1971, Philip Zimbardo assembled the infamous Stanford Prison Experiment, in which 24 college students were randomly assigned guard or prisoner roles in a simulated prison. Within 36 hours, guards became dominant, creating punishments and restrictions, while prisoners became silent and despairing. Zimbardo’s team had to shut down the two-week experiment after only six days.6 While critics later questioned the methodology and ethics of the study, the core finding remained: structures—uniforms, daily routines, hierarchical setups—can pressure participants into abusive obedience without explicit orders.

The late 1980s brought a subtler turn with Willem Meeus and Quinten Raaijmakers in the Netherlands. Participants were asked to shame a job applicant over a phone call, and told it was part of official psychological testing. With no visible authority figure, most still complied, aping humiliating commands.3 This study demonstrated how obedience can distort workplace norms, and how quiet coercion, not shocks or force, drives the social mission of institutions.

By the early 2000s, psychologists began to wonder: had society changed enough to resist authority more than in Milgram’s era? Jerry M. Burger, a social psychologist, revisited that question in a careful modern replication of Milgram’s experiment.7 His findings were sobering. Despite decades of cultural shifts and growing public awareness of the original experiments, people still obeyed the experimenter. In fact, the rates of compliance were remarkably close to those Milgram recorded nearly fifty years earlier. Even when participants watched someone else refuse, they rarely followed suit. The presence of authority—polite, composed, and institutional—remained a powerful force. Burger’s study confirmed what many feared: our instinct to obey hasn’t faded. It still overrides personal hesitation, moral discomfort, and even social modeling. The study didn’t just replicate a procedure; it echoed a troubling truth about human behavior that transcends time.

Technological shifts introduced algorithmic authority. Modern studies, such as one in 2023, revealed that people often follow AI prompts with near-identical obedience levels as human orders.8 Even when participants know the prompt arises from a computer, they adjust settings or follow suggestions they might not trust, citing AI’s institutional legitimacy. Authority has evolved from ancestral rulers to digital systems.9

Today, obedience studies focus on resistance training, safety culture, and ethical systems. In aviation, cockpit resource management taught junior crew to speak up and challenge authority.10 Hospitals instituted "stop-the-line" protocols so nurses can demand clarification on instructions.11 These are real-world experiments in flattening authority gradients—the invisible but powerful hierarchies that make it hard for subordinates to question those in charge, even when they notice something wrong.

Obedience shapes modern society, driving everything from vaccine mandates to managerial compliance. In health settings, hierarchical pressures can suppress whistleblowing, causing tragic outcomes like infections or preventable deaths. In military systems, unchecked obedience led to atrocities, spurring reforms around rules of engagement and legal review boards. In corporate finance, unchecked obedience contributed to fraud and systemic failure. These lessons in tech, medicine, law, and education show obedience exists not to judge individuals, but to shape structures that enable moral courage.

That’s the haunting legacy of obedience research: it reveals a quiet power at the heart of systems. Its lessons resonate beyond shock machines or surveillance labs. Understanding its psychological fabric helps design social, institutional, and technological structures that balance coordination with conscience, supporting obedience when it's necessary, dissent when it’s moral, and ethical judgment when it counts.

People

Stanley Milgram

An American social psychologist who revolutionized the field with his 1961 obedience experiments at Yale University. Provoked by the trial of Nazi officer Adolf Eichmann, Milgram sought to understand how ordinary people could commit harmful acts under authority. His experimental setup, involving fake electric shocks administered by participants, revealed an unsettling truth: 65% of people would follow orders to the point of extreme harm. Milgram’s studies ignited ethical debates in research and changed how psychologists interpret human behavior in authoritarian contexts.

Charles K. Hofling

A psychiatrist by training, Charles Hofling pushed obedience research out of the lab and into real-world clinical settings. In a now-famous 1966 study, he tested how nurses would respond to a fake doctor’s orders to administer a dangerous drug, and nearly all complied. This simple field experiment exposed a silent crisis in healthcare: the dangerous power of perceived authority over professional judgment. Hofling’s findings remain central in hospital training programs aimed at encouraging ethical resistance. His work demonstrated that hierarchy, not just intention, can be a dangerous catalyst.5

Philip Zimbardo

Philip Zimbardo, a psychologist at Stanford University, took Milgram’s insights into the realm of role-playing and institutional power. His 1971 Stanford Prison Experiment simulated a prison environment with volunteers as guards and prisoners, revealing how quickly people conform to oppressive roles. Zimbardo was shocked by how deeply the “guards” embraced cruelty, and how passively the “prisoners” accepted abuse. Though ethically controversial, the study offered enduring insights into the mechanisms behind obedience, role identity, and deindividuation.6

Willem Meeus & Quinten Raaijmakers

Willem Meeus and Quinten Raaijmakers expanded the concept of obedience beyond physical harm into the psychological realm. In their 1986 study, participants were instructed to make hostile remarks toward job applicants during a fake interview, believing it was part of a psychological assessment. Most complied, even when visibly distressed, revealing how bureaucratic systems mask coercion behind polite procedure. Their research coined the term “administrative obedience,” shifting focus from dramatic to everyday forms of compliance. Today, their findings help inform workplace ethics and the dangers of institutional pressure.3

Jerry M. Burger

A social psychologist best known for his ethically updated replication of Milgram’s obedience experiment in 2009. A professor at Santa Clara University, Burger was curious whether modern participants would still follow harmful orders from an authority figure. His study revealed an unsettling continuity: obedience levels remained strikingly close to Milgram’s, even with added ethical safeguards. Burger’s work reignited public interest in obedience, highlighting how deeply situational pressures can override individual conscience.7 

behavior change 101

Start your behavior change journey at the right place

Impacts

Obedience is a loaded word. We think of rule-followers and rule-breakers, of military orders or classroom instructions. But beneath that binary lies a deeper truth: obedience isn’t about submission—it’s about structure. It’s a social tool. And when harnessed well, it becomes a catalyst for learning, safety, and innovation. Across education, healthcare, and technology, modern thinkers are re-engineering obedience, not to suppress autonomy, but to amplify it.

When silence sparks thinking

In a noisy classroom filled with fast-paced questions and nervous students, silence doesn’t feel natural. But in the 1970s, Mary Budd Rowe strapped timers to her wrists and visited public school science classrooms to measure exactly that: silence. She found that the average teacher waited less than a second after asking a question before jumping in or moving on.12 That speed didn’t give students time to think, let alone feel safe enough to respond thoughtfully.

Rowe then coached teachers to stretch their wait time after asking a question, from the usual one second to a full three to five. It seemed small. But when she recorded and analyzed classroom interactions, the shift was dramatic. Students started reasoning, connecting ideas, and using their own words. One fifth-grader, who had barely whispered “photosynthesis,” went on to describe it like a factory making food.13 The silence had given him space to think. What’s often missed is what changed in the teachers. With longer pauses, they stopped interrupting. They asked more open-ended questions. They listened. In that space, the power dynamic softened. Obedience was about structure. By holding the pause, teachers reshaped how authority worked in the room. They weren’t demanding answers; they were creating conditions that encouraged thoughtful engagement. The result was a more collaborative kind of obedience: students followed the rules, but also took intellectual risks.

Today, digital learning platforms are putting this insight to work. In a study led by Dunlosky and colleagues, students were asked to explain their reasoning before seeing the correct answer.14 Those who complied with this prompt retained more information a week later than those who didn’t. The instruction was simple: Write out your thought process before continuing. This was a low-stakes act of obedience, no pressure, no punishment. But when students followed the instruction, the results were dramatic. That moment of compliance forced reflection. It interrupted passive clicking and triggered deeper processing. Like Rowe’s classroom pause, this digital prompt wasn’t about controlling behavior; it structured it, nudging learners to engage more actively. Obedience here opened the door to better thinking.

In essence, we’ve moved from obedience as silence to obedience as scaffolding. Students are given space to build their own thoughts within structured prompts. That’s how obedience becomes a launchpad, not a leash.

Whistle before you cut

Picture this: it’s 3:00 a.m. in a quiet hospital ward. A nurse receives a call from someone claiming to be a physician. They instruct her to administer a drug, “Astroten,” at 20 mg. The bottle says 10 mg is the max dose, and the drug isn’t even listed in the hospital formulary. The doctor is unknown. The order violates three clear hospital policies.

Yet in Charles Hofling’s chilling 1966 field study, 21 nurses faced that exact situation, and 19 began to carry out the order.5 The researchers never let the injection happen, but the setup was real. Nurses believed they were following protocol. What mattered most wasn’t training or rational risk; it was the power of perceived authority.

That experiment prompted an overhaul in nursing ethics and communication norms. More importantly, it launched a new wave of intervention research: how can we train obedience to authority without compromising patient safety? The answer lies in empowering structured disobedience. Many hospitals now use protocols borrowed from manufacturing, where any team member can halt a procedure to flag a concern.11 In high-pressure environments, these pause moments work like Rowe’s wait-time: they slow things down, give people breathing room, and shift authority from hierarchy to shared accountability.

Take the WHO’s Surgical Safety Checklist: it begins with every team member, from lead surgeon to intern, introducing themselves by name. Then, before the first incision, a full confirmation of the procedure, patient, site, and expected complications is conducted. In a massive multicenter trial across eight countries, this checklist cut serious surgical complications by 36% and slashed mortality by nearly half.15 That’s not just protocol; it’s principled obedience, engineered to include everyone.

The effects last beyond the operating room. Follow-up studies show that institutions that use these checklists also report higher levels of psychological safety, better teamwork, and improved reporting of near misses.16 Obedience, in this world, isn’t about keeping quiet. It’s about making sure the right voices get heard at the right time.

The algorithm whisperer

Imagine being immersed in a virtual world where you’re the one delivering shocks, not in a real lab, but through a headset, clicking buttons on a digital console. In a study at University College London, González-Franco and colleagues recreated Milgram’s famous obedience experiment inside virtual reality.17 This wasn’t the same as Burger’s partial replication; it was a new digital adaptation. While the setup echoed Milgram’s original design, key elements shifted. The shocks weren’t real—the learner was a virtual avatar, and the authority figure appeared in the simulation. What remained was the pressure: a voice urging participants to continue, an escalating series of “shocks,” and the visible distress of the learner. The goal wasn’t to copy every variable, but to explore how obedience plays out when the environment is artificial and the consequences feel simulated.

Participants were placed in a virtual simulation and told they were helping with a learning experiment. Their job was to quiz a digital avatar, designed to look and move like a realistic human, on word pair memorization. Every time the avatar gave a wrong answer, participants were instructed to administer an electric “shock” by clicking a button. With each mistake, the shock level increased, and the avatar reacted with escalating pain: flinching, grimacing, even begging to stop. Most participants followed the commands, echoing Milgram’s original results, but the researchers were watching more than compliance. They tracked heart rate spikes, pauses before each click, and quiet signs of discomfort. Some participants whispered “sorry” before delivering shocks. Others subtly coached the avatar to avoid punishment. Obedience was still there, but layered with resistance, hesitation, and empathy, revealing how people can follow orders while quietly pushing back.

These findings sparked conversation among tech designers. If users will obey even simulated authority figures under pressure, what happens when AI becomes that figure? The response: design for pushback. In finance, several firms introduced override checkpoints in AI loan-approval systems. Advisors were required to enter a justification any time they overrode or accepted a machine-generated suggestion. One firm saw a decrease in risky approvals and a rise in staff confidence in their final decisions.18 Obedience became conscious.

On the global stage, the EU’s AI Act now mandates human-in-the-loop oversight for all high-risk systems, including those used in medicine, criminal justice, and employment decisions.19 It’s no longer enough to obey what a machine suggests. The system must ask, “Are you sure?” and wait for the human to think. This isn’t just tech regulation; it’s a moral model. Designers now engineer systems to expect doubt, to embed friction, to reward hesitation. Because in a world filled with algorithmic influence, the most responsible obedience might be the kind that resists just enough.

Controversies

Obedience isn’t a simple matter of following orders—it’s a catalyst for major ethical, institutional, and technological debates. As we explore three pivotal controversies, it becomes clear that obedience has the power both to enlighten and to blind. Each of these debates forces us to reckon with when obedience supports human well-being and when it slips into moral danger, structural compulsion, or algorithmic overreliance.

Is psychological insight worth the emotional cost?

In the wake of World War II, Stanley Milgram sought to probe a haunting question: how could everyday people become agents of harm simply because someone in authority told them to? While his 1963 shock experiment became a cultural flashpoint, Milgram’s lesser-known “Lost Letter” study offered a quieter but equally chilling lens on obedience.20 In this field experiment, he “lost” stamped, addressed letters in public spaces, some intended for clearly sympathetic recipients (like medical research foundations), others for controversial groups (such as fictitious organizations promoting communist or neo-Nazi causes). What he found was telling: people were far more likely to mail letters to recipients aligned with mainstream values than to those representing stigmatized ideologies.

This wasn’t an overt obedience test, but it revealed the subtle force of social conformity and complicity. People weren’t pressured or commanded; they acted freely. Yet their behavior still mirrored a quiet obedience to dominant cultural norms. Even when anonymity was guaranteed, many chose not to help when the target of aid was ideologically threatening. Milgram argued that this revealed the infrastructure of obedience not just in explicit commands but in the soft contours of social approval.21 Whether in returning letters or pressing buttons, his studies suggested that people conform when the social frame makes dissent feel wrong, awkward, or dangerous. 

Others, however, found Milgram’s methods deeply troubling. Diana Baumrind, a leading developmental psychologist, criticized the study for crossing ethical boundaries. Participants were led to believe they were inflicting real harm, and many showed signs of acute distress—twitching, stuttering, and, in some cases, nearing emotional collapse.22 She warned that the emotional trauma could linger long after the lab session ended. This concern led to the concept of inflicted insight, the idea that participants might be psychologically harmed by discovering their own willingness to obey harmful commands.23 But ethics weren’t the only controversy. Critics also questioned the study’s design. Some argued that participants may have guessed the shocks weren’t real, casting doubt on the results. Others pointed to selection bias: most participants were American men, limiting how broadly the findings could apply. 

These concerns reshaped research practices. Institutional Review Boards now mandate fully informed consent, immediate debriefing, and the option to withdraw without penalty, all designed to minimize emotional risk.24 Critics argue that these protective layers may dampen a study's capacity to fully reveal obedience's depths. Without the emotional intensity Milgram induced, can obedience research still expose its darkest dimensions? It’s a tension between safeguarding participants and preserving experimental validity, a debate that defines the boundary between ethical care and scientific exploration.

Can bad systems corrupt good people or just the willing?

In 1971, psychologist Philip Zimbardo launched what would become one of the most disturbing and iconic studies in social psychology: the Stanford Prison Experiment. But it wasn’t his first time exploring how authority and anonymity distort behavior.6 Years earlier, Zimbardo had been deeply engaged with the concept of deindividuation, the idea that people are more likely to lose moral inhibition and act impulsively when their personal identity is obscured.

In one of his foundational studies, participants dressed in identical robes and hoods, while others were blindfolded and stripped of social cues.25 The results were unsettling: when people couldn’t be individually recognized, they became far more likely to deliver punishments, conform to group behavior, and abandon personal accountability. Zimbardo argued that anonymity, combined with a group-defined role, dissolved the internal moral compass. His early research laid the groundwork for a chilling question: What happens when everyday people are given power without oversight?

Yet the Stanford Prison Experiment has faced substantial criticism. Thomas Carnahan and Sam McFarland (2007) conducted personality assessments and found that participants exhibited higher levels of aggression, authoritarianism, and dispositional empathy issues, suggesting they were predisposed to dominant behaviors.26 Thibault Le Texier (2019) dug into archives and interviews, uncovering evidence that Zimbardo’s team had coached guards through orientation sessions that included implicit or explicit prompts to enforce control.27 This suggests the guards may have been following perceived behavioral scripts, rather than autonomously succumbing to institutional pressure.

The implications of this critique are profound. If structural environments can override individual morals, then institutions, prisons, militaries, and corporations must design systems that protect against passive compliance. Training should reinforce personal moral responsibility and create oversight to prevent blind obedience. Conversely, if predisposition and researcher influence explain abusive behavior, then reform should focus on recruiting individuals with a strong moral compass and leadership training. This debate shapes everything from corporate compliance programs to the legal acceptance of “just following orders” as a defense, challenging how societies balance collective and personal responsibility.

Can we teach AI to trigger critical thinking?

At Cornell University, Julian Senoner and his team explored this by replacing typical “black box” predictions with interactive heatmaps. These visual tools highlighted what the AI model was focusing on, be it a shadow in a lung x-ray or a flaw in a manufactured part. In randomized trials with factory inspectors and radiologists, those using the explainable AI system made five times fewer diagnostic errors.28 It wasn’t just accuracy that improved; participants described feeling more in the loop. They said the tool didn’t dictate, but collaborated. That shift—AI as a teammate, not a commander—transformed how people interacted with the machine. They clicked, zoomed, and compared multiple perspectives before locking in a choice. Obedience became informed engagement.

Still, these friction-based interventions raise difficult trade-offs. In high-stakes environments like hospital trauma bays or wildfire containment systems, even a five-second delay could prove fatal. Critics warn that introducing too much interaction might stall action or overwhelm users with cognitive clutter. Others, like Parasuraman and Manzey, have documented the risk of automation complacency, a subtle form of moral abdication where users assume they’ve done enough just by following the system’s reflective steps.29 If pausing to click a box becomes a stand-in for deeper analysis, we haven’t eliminated blind obedience. We’ve just disguised it.

Governments are wrestling with this tension. The EU’s 2024 AI Act mandates human oversight for high-risk systems, requiring that users be able to interrupt or override automated recommendations.19 But the law offers no clear direction on how much friction is enough, or what kinds of prompts genuinely foster critical thinking. That puts the burden on designers. Do they create smooth, seamless AI workflows that minimize friction? Or deliberately engineer “speed bumps” that invite questions and slow compliance? Scholars like Andreas Holzinger argue we must do both. The goal isn’t just interpretability; it’s usability. AI must be transparent and responsive, nudging us toward better decisions without tripping us up in the process.30 As we hand more power to algorithms, the stakes grow higher. Thoughtful design can make obedience to AI safer and smarter, but only if that obedience demands something back: reflection.

Case Studies

How a simple classroom game rewrote childhood and beyond

Picture a bustling first-grade classroom in 1985 Baltimore: chairs clatter, crayons roll, and teachers juggle instruction while keeping an eye on wandering hands. That’s where psychiatrist Sheppard Kellam stepped in, armed not with discipline charts or detentions, but with an idea: what if children could learn self-control by cooperating, rather than being singled out?

In 41 classrooms across 19 urban schools, Kellam and his colleagues introduced the Good Behavior Game (GBG). Each classroom divided students into roughly even teams, blending kids with different temperaments: some chatty, some zoned out, some full of energy. The rules set clear classroom expectations: raise your hand before speaking, stay seated, follow the lesson quietly, and listen when others talk. Each team started with a clean slate.31

What unfolded was a ritual of collective accountability. Three times a week, during specially designated “game periods” lasting 10 to 45 minutes, teachers tracked behavior using simple tally sheets. Every time a child broke a rule—shouting, interrupting, leaving their seat—their team earned a mark. If a team stayed below a threshold of infractions, they won a reward: extra recess, small prizes, or spirited praise from the teacher. These weren’t punishments aimed at individuals; they were gentle nudges prompting the group to self-regulate.

The researchers didn’t just rely on teacher reports. They embedded observers in classrooms, timing how long students stayed focused, counting interruptions, and manually logging each infraction and reward. They kept notes on eye contact, peer reminders (“Johnny, sit down!”), and the atmosphere when teams lost or won, a mix of pride and playful competition. Within a week, the change was dramatic. Noise levels dropped. Teachers reported fewer verbal interventions needed, shifting from reactively tuning out chaos to proactive engagement with lessons. Children started reminding their peers to follow the rules: one kid pointing out “You forgot to raise your hand!” as if he’d just climbed the honor roll. 

The long-term follow-up on these children is what truly amazes. Nearly 20 years later, when Kellam’s participants turned 19–21, researchers tracked them down through high school alumni records, mailers, and social services. Young adults from GBG classrooms, particularly boys deemed “aggressive” in childhood, showed significantly lower instances of substance abuse, daily smoking, and antisocial behavior1. One participant reflected, decades on, “I learned to watch myself, not break the line, even when no teacher was watching.”

This classroom experiment wasn’t just about obedience; it was about cooperative self-governance. It rewrote the message from “Follow orders or be punished” to “Team success comes from shared self-control.” It planted small seeds of responsibility and social awareness in young hearts, and those seeds sprouted for decades. Even better, Kellam’s dataset didn't just stop at behavior. It included psychiatric diagnoses, drug use surveys, school transcripts, arrest records, and even suicide risk screenings. The correlation between early cooperative obedience and later life outcomes was robust: GBG grads were less likely to land in juvenile detention, to rely on mental health services, or to suffer early substance dependence.

What makes this story compelling is how it tells a story, not just presents a chart. This is about a little boy whose team rallied him to put down a rogue toy, a teenager who remembered his childhood lesson to "manage your space," and most profoundly, a researcher who realized that meaningful obedience can be part of a playful community, not just punishment. It’s a radical shift from seeing obedience as submission to seeing it as shared purpose.

The like that launched a thousand opinions

It started with a single click. In 2013, researchers Lev Muchnik, Sinan Aral, and Sean Taylor slipped into the backend of an unnamed social news aggregation website, widely believed to be similar in structure to Reddit. Their goal? Test whether a crowd could be steered by a whisper. Specifically, could just one vote, an upvote or downvote, alter the way thousands of strangers judged public comments? They didn’t just theorize. They took action.32

The team quietly assigned 101,281 newly posted comments into three invisible categories: up-treated (which got a fake +1 rating), down-treated (a fake -1), and control (left untouched). These tweaks happened automatically, the moment each comment was published. No alerts. No hints. Just an imperceptible tilt, a digital coin flip at the moment of creation. Then, the comments were left to fend for themselves in the wild.

Over the next five months, more than 10 million people viewed and voted on these comments. None of them knew they were participants in a social influence experiment. They were just users, scrolling through comment sections, rating what they read. But behind the scenes, Muchnik and his colleagues were watching a behavioral domino effect unfold.

The results were staggering. A single artificial upvote significantly boosted a comment’s odds of receiving real upvotes later. That first viewer was 32% more likely to give the comment a thumbs-up compared to a comment with no manipulation1. And the ripple didn’t stop there. Over time, those early upvotes snowballed; comments that received them ended up with final scores that were 25% higher on average than those in the control group. This wasn’t just correlation. It was causality in action, engineered through randomized assignment.

But what about negative signals? Down-treated comments initially saw more downvotes, yes. But then came something unexpected: the crowd pushed back. Many users seemed to notice the unfairness, subconsciously or not, and responded with corrective upvotes. In fact, negatively manipulated comments ended up with final scores statistically indistinguishable from the control group. In short, people corrected what looked like undeserved negativity. But when it came to false positivity? That got rewarded and magnified.

Even more compelling were the hidden layers of nuance. The researchers broke the data down by topic and discovered that the size of these ripple effects varied. Political discussions showed the strongest herding. Comments in this category, once upvoted, became lightning rods for more agreement. On less contentious topics like general news or entertainment, the same manipulation had a softer echo. Obedience, it seemed, depended not just on the signal, but on the stakes.

There was another twist: social context mattered. On this platform, users could label each other as “friends” or “enemies.” Friends were far more likely to follow the crowd and amplify positive ratings. Enemies, on the other hand, seemed largely immune. They neither followed nor corrected, like neutral islands in a sea of social sway. This added a layer of relational psychology to what initially looked like a simple rating game.

Ultimately, what the study exposed was a hidden architecture of influence. In theory, rating systems exist to capture the public’s voice. But in practice, they often create that voice. One arbitrary vote could bend perception. One quiet cue could rewrite consensus. There were no rules. No enforcement. No pressure. Just an interface that nudged people to agree, and a population that unknowingly obeyed.

Related TDL Content

Authority Bias

In this article, TDL unpacks authority bias, the powerful tendency to accept opinions or instructions from figures we consider authoritative. This bias lies at the heart of obedience, helping explain why Milgram’s participants continued to administer shocks, and why nurses in Hofling’s experiment obeyed questionable medical orders. By examining real-world examples (like trusting medical experts without checking evidence), the article shows how authority can bypass our critical thinking. 

Conformity

TDL’s deep dive into conformity explores why we often “go with the flow,” even when we know better. Linking classic Asch-style line-judgment studies to modern examples, like peer pressure in workplaces and online communities, it reveals how social influence often trumps independent thought. Readers will learn about the difference between informational conformity (accepting information because others seem informed) and normative conformity (wanting to fit in).

References

  1. Milgram, S. (1963). Behavioral study of obedience. Journal of Abnormal and Social Psychology, 67(4), 371–378. https://doi.org/10.1037/h0040525
  2. Blass, T. (1999). The Milgram paradigm after 35 years: Some things we now know about obedience to authority. Journal of Applied Social Psychology, 29(5), 955–978. https://doi.org/10.1111/j.1559-1816.1999.tb00134.x
  3. Meeus, W. H., & Raaijmakers, Q. A. (1986). Administrative obedience: Carrying out orders to use psychological-administrative violence. European Journal of Social Psychology, 16(4), 311–324. https://doi.org/10.1002/ejsp.2420160402
  4. Arendt, H. (1963). Eichmann in Jerusalem: A report on the banality of evil. Viking Press.
  5. Hofling, C. K., Brotzman, E., Dalrymple, S., Graves, N., & Pierce, C. M. (1966). An experimental study in nurse-physician relationships. The Journal of nervous and mental disease, 143(2), 171–180. https://doi.org/10.1097/00005053-196608000-00008
  6. Zimbardo, P. G. (2007). The Lucifer effect: Understanding how good people turn evil. Random House.
  7. Burger J. M. (2009). Replicating Milgram: Would people still obey today?. The American psychologist, 64(1), 1–11. https://doi.org/10.1037/a0010932
  8. Walker, R., Dillard-Wright, J., & Iradukunda, F. (2023). Algorithmic bias in artificial intelligence is a problem-And the root issue is power. Nursing outlook, 71(5), 102023. https://doi.org/10.1016/j.outlook.2023.102023
  9. Green, J., & Williams, S. (2018). Digital compliance: AI voice assistants and authority. Psychology & Technology, 2(1), 45–58. https://doi.org/10.1037/pta0000108
  10. Fidan, C., & Demirarslan, Y. (2024). The effects of crew resource management on flight safety culture. The Aeronautical Journal, 128(1326), 1743–1766. https://doi.org/10.1017/aer.2023.113
  11. Rotter, T., Plishka, C. T., Adegboyega, L., Fiander, M., Harrison, E. L., Flynn, R., Chan, J. G., & Kinsman, L. (2017). Lean management in health care: effects on patient outcomes, professional practice, and healthcare systems. The Cochrane Database of Systematic Reviews, 2017(11), CD012831. https://doi.org/10.1002/14651858.CD012831
  12. Rowe, M. B. (1972). Wait‑time and rewards as instructional variables: Their influence on language, logic, and fate control. Journal of Research in Science Teaching, 11(2), 81–94. https://doi.org/10.1002/tea.3660110202
  13. Rowe, M. B. (1986). Wait time: Slowing down may be a way of speeding up!. Journal of Teacher Education, 37(1), 43–50. https://doi.org/10.1177/002248718603700110
  14. Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). Improving students’ learning with effective learning techniques: Promising directions from cognitive and educational psychology. Psychological Science in the Public Interest, 14(1), 4–58. https://doi.org/10.1177/1529100612453266
  15. Haynes, A. B., et al. (2009). A surgical safety checklist to reduce morbidity and mortality in a global population. New England Journal of Medicine, 360(5), 491–499. https://doi.org/10.1056/NEJMsa0810119
  16. Urbach, D. R., et al. (2015). Effect of the World Health Organization checklist on surgical outcomes: A prospective, cluster randomized trial. Journal of Gastrointestinal Surgery, 19(5), 935–942. https://doi.org/10.1007/s11605-015-2772-9
  17. González‑Franco, M., Slater, M., Birney, M. E., Swapp, D., Haslam, S. A., & Reicher, S. D. (2018). Participant concerns for the Learner in a virtual reality replication of the Milgram obedience study. PLoS ONE, 13(12), e0209704. https://doi.org/10.1371/journal.pone.0209704
  18. Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To Trust or to Think: Cognitive forcing functions can reduce overreliance on AI in AI‑assisted decision‑making. Proceedings of the ACM on Human‑Computer Interaction, 5(CSCW1), Article 188. https://doi.org/10.1145/3449287
  19. Fink, M. (2025). Human oversight under Article 14 of the EU AI Act. Computer Law & Security Review. Advance online publication. https://doi.org/10.1080/17579961.2023.2245683
  20. Milgram, S., Mann, L., & Harter, S. (1965). The lost-letter technique: A tool of social research. Public Opinion Quarterly, 29(3), 437–438. 
  21. Milgram, S. (1974). Obedience to authority: An experimental view. New York: Harper & Row.
  22. Baumrind, D. (1964). Some thoughts on ethics of research: After reading Milgram’s “Behavioral Study of Obedience.” American Psychologist, 19(6), 421–423. https://doi.org/10.1037/h0040128
  23. Harnett J. D. (2021). Research Ethics for Clinical Researchers. Methods in molecular biology (Clifton, N.J.), 2249, 53–64. https://doi.org/10.1007/978-1-0716-1138-8_4
  24. Levine, R. (2018). Ethics and regulation of clinical research (3rd ed.). Yale University Press.
  25. Zimbardo, P. G. (1969). The human choice: Individuation, reason, and order versus deindividuation, impulse, and chaos. In W. J. Arnold & D. Levine (Eds.), Nebraska Symposium on Motivation (Vol. 17, pp. 237–307). Lincoln: University of Nebraska Press.
  26. Carnahan, T., & McFarland, S. (2007). Revisiting the Stanford Prison Experiment: Could participant self-selection have led to the cruelty? Personality and Social Psychology Bulletin, 33(5), 603–614. https://doi.org/10.1177/0146167206292689
  27. Le Texier, T. (2019). Debunking the Stanford Prison Experiment: A reexamination of archival evidence and participant interviews. American Psychologist, 74(7), 823–839. https://doi.org/10.1037/amp0000401
  28. Senoner, J., Schallmoser, S., Kratzwald, B., Feuerriegel, S., & Netland, T. (2024). Explainable AI improves task performance in human–AI collaboration. ArXiv. https://doi.org/10.48550/arXiv.2406.08271
  29. Parasuraman, R., & Manzey, D. (2010). Complacency and bias in human use of automation: An attentional integration framework. Human Factors, 52(3), 381–410. https://doi.org/10.1177/0018720810376055
  30. Holzinger, A., Carrington, A., & Müller, H. (2019). Interactive machine learning: Experimental evidence for the human in the algorithmic loop. Applied Intelligence, 49(7), 2401–2414. https://doi.org/10.1007/s10489-018-1361-5
  31. Kellam, S. G., Brown, C. H., Poduska, J. M., Ialongo, N. S., Wang, W., Toyinbo, P., Petras, H., Ford, C., Windham, A., & Wilcox, H. C. (2008). Effects of a universal classroom behavior management program in first and second grades on young adult behavioral, psychiatric, and social outcomes. Drug and alcohol dependence, 95 Suppl 1(Suppl 1), S5–S28. https://doi.org/10.1016/j.drugalcdep.2008.01.004
  32. Muchnik, L., Aral, S., & Taylor, S. J. (2013). Social influence bias: A randomized experiment. Science, 341(6146), 647–651. https://doi.org/10.1126/science.1240466

About the Author

White guy wearing a white lab coat over a baby blue dress shirt.

Adam Boros

Researcher, Mount Sinai Hospital

Adam studied at the University of Toronto, Faculty of Medicine for his MSc and PhD in Developmental Physiology, complemented by an Honours BSc specializing in Biomedical Research from Queen's University. His extensive clinical and research background in women’s health at Mount Sinai Hospital includes significant contributions to initiatives to improve patient comfort, mental health outcomes, and cognitive care. His work has focused on understanding physiological responses and developing practical, patient-centered approaches to enhance well-being. When Adam isn’t working, you can find him playing jazz piano or cooking something adventurous in the kitchen.

About us

We are the leading applied research & innovation consultancy

Our insights are leveraged by the most ambitious organizations

Image

I was blown away with their application and translation of behavioral science into practice. They took a very complex ecosystem and created a series of interventions using an innovative mix of the latest research and creative client co-creation. I was so impressed at the final product they created, which was hugely comprehensive despite the large scope of the client being of the world's most far-reaching and best known consumer brands. I'm excited to see what we can create together in the future.

Heather McKee

BEHAVIORAL SCIENTIST

GLOBAL COFFEEHOUSE CHAIN PROJECT

OUR CLIENT SUCCESS

$0M

Annual Revenue Increase

By launching a behavioral science practice at the core of the organization, we helped one of the largest insurers in North America realize $30M increase in annual revenue.

0%

Increase in Monthly Users

By redesigning North America's first national digital platform for mental health, we achieved a 52% lift in monthly users and an 83% improvement on clinical assessment.

0%

Reduction In Design Time

By designing a new process and getting buy-in from the C-Suite team, we helped one of the largest smartphone manufacturers in the world reduce software design time by 75%.

0%

Reduction in Client Drop-Off

By implementing targeted nudges based on proactive interventions, we reduced drop-off rates for 450,000 clients belonging to USA's oldest debt consolidation organizations by 46%

Read Next

Notes illustration

Eager to learn about how behavioral science can help your organization?