Human-In-The-Loop Systems

What are Human-in-the-loop Systems?

Human-in-the-loop (HITL) systems integrate human oversight, feedback, and decision-making into artificial intelligence (AI) or machine learning processes. These systems rely on humans to label data, verify outputs, correct errors, and guide AI learning, improving accuracy, fairness, and reliability in applications ranging from healthcare and fraud detection to voice assistants and autonomous systems.

The Basic Idea

Many of us take advantage of the cruise control feature in our cars when we’re driving at a consistent pace on an empty highway. Newer car models also have an adaptive cruise control system, which uses sensors and AI to adjust the car’s speed to maintain a safe distance from other vehicles. If your cruise control is set to 80 miles per hour, but a car merges onto your lane, the sensor will reduce speed until the path is clear.

However, imagine that you’re driving along, using adaptive cruise control, when you suddenly begin to slow down unexpectedly. You look around and see no cars in sight; however, you notice a plastic bag up ahead. The sensor has mistaken the plastic bag for a car. Confident there is no danger, you push your foot on the gas pedal to get the car to speed up again.

In this instance, you’ve acted like a “human-in-the-loop” of the adaptive cruise control system. You intervened when the system was not functioning properly. HITL systems are collaborative approaches where human input and expertise are integrated throughout the lifecycle of a machine learning or AI system to provide guidance and feedback that help the system improve.1 

HITL systems are important for tasks and areas that involve judgment, contextual understanding, and uncertainty. The feedback that a human provides allows AI models to adapt and evolve based on real-world scenarios, and become more personalized to user preferences.

“

The thing that always works the best is the intelligence from the human and the raw power from the machine. That’s what we need. We don’t replace the human. We amplify human intelligence. We augment the human.


— Gary King, Professor of Government at Harvard University and Director of the Institute for Quantitative Social Science (IQSS).2

Key Terms

Artificial Intelligence: Tools and algorithms designed to train computers and machines to process and analyze data in a way that mimics the learning, reasoning, and problem-solving capabilities of humans. AI is used to train systems to perform tasks that historically would have required human cognitive skills. 

Machine Learning: A branch of AI, focused on developing computers and machines that can learn from data to adapt and improve over time, mimicking the way that humans learn. Instead of being explicitly programmed, machines learn over time through the analysis of vast amounts of data, which allows them to identify patterns and generate predictions and recommendations. 

Training Data: Information used to teach a machine learning model how to analyze data, recognize patterns, and make decisions. As machine learning does not involve the explicit programming of an algorithm, training data is required for it to develop its own algorithms. Often, the training data’s attributes are labeled or corrected by a human as a form of HITL, to contextualize the data for the AI system. 

AI Feedback Loops: The process through which an AI system receives feedback, either via additional data or human input, on its performance so it can adapt and improve. For example, after using a customer service chatbot, you may be asked to complete a survey to indicate how well the tool resolved your issues. Through machine learning, your feedback is incorporated into the algorithm to make adjustments to the system and improve performance. 

Quality Assurance: The process through which we ensure that an output, whether a product or service, meets quality standards. In HITL systems, quality assurance can be in the form of human checks on AI-generated outputs to catch errors and intervene when necessary. 

History

In 1960, American psychologist and pioneer in human-computer interaction, J.C.R. Licklider, published a paper titled “Man-Computer Symbiosis,” in which he outlined his belief that eventually, humans and computers would cooperate very closely to enhance decision-making. He believed that humans and computers would form a symbiotic relationship, a close and long-term interaction between two different species, without the need for a predetermined program. This was a radical idea at the time, as computers were seen purely as tools rather than as collaborators in a team setting. Licklider envisioned computers taking over routine tasks and making them more efficient, with an intuitive capability that meant they didn’t need to have a predetermined program.3 

Following Licklider’s paper, and the emergence of personal computing, psychologists began to take an interest in the relationship between machines and humans, and in the early 1980s, human-computer interaction became its own field. Now that people beyond information technology professionals were using computers, usability became a focus area and psychologists wanted better to understand the emotional and social aspects of these interactions. 

In 1983, Danish engineer and psychologist Jens Rasmussen researched how humans and automation systems interacted—particularly in complex, safety-critical systems like in aviation or energy. He recognized that people often blamed human error when mistakes happened in these partly automated systems. To break down what kind of cognitive processes were occurring in these instances of human error, he developed the Skill, Rule, Knowledge (SRK) framework. The framework outlined three levels of cognitive control:

  • Skill: low-level subconscious or motor cognition, like pushing the wrong button because it is close to another one
  • Rule: mid-level pattern recognition, where a mistake occurs while following learned rules, such as following a standardized checklist in a more complex situation, like an edge case, where the rules don’t apply
  • Knowledge: high-level environmental, social, and emotional factors that lead to errors, such as not informing superiors about an issue because the person fears they will be blamed.4

Rasmussen suggested that machines could handle skill and routine tasks, but humans need to step in at the knowledge level. This model helped describe how humans make decisions in complex environments, demonstrating where humans need to be placed in the loop of safety systems. Throughout the 1980s and 1990s, this model of semi-automation with human oversight was applied in aviation, space exploration, and military and defense simulations.4

HITL systems expanded these fields in the early 2000s through open-source projects, where software source code was freely available for users to modify and share. One of the first open-source computer operating systems was Linux, developed by Finnish software engineer Linus Torvalds. This laid the groundwork for crowdsourcing the public to complete microtasks that are difficult for computers but easy for humans. The first large-scale crowdsourcing platform for these aptly named “human intelligence tasks” was launched in 2005 by Amazon: Amazon Mechanical Turk. An example of such a task could be, “Is there a burger joint in this photograph?”—allowing humans to label the data to help the computer better learn how to identify and recognize objects in images.5

These developments marked a turning point, moving HITL models from specialized labs into the hands of the public, leading to accelerated AI research and a redefinition of the relationship between humans and technology.

People

J.C.R. Licklider

An American psychologist and computer scientist whose most notable contribution to the field was laying the groundwork for computer networking—the ability for computers to connect to communicate data electronically—and the predecessor of the Internet, Advanced Research Projects Agency Network (ARPANET). Licklider’s interest in the relationship between computers and humans began when he was teaching at the Massachusetts Institute of Technology (MIT) in 1950, where he researched people’s interaction with a pilot computerized air defense system. This led to his belief, as outlined in his 1960 paper “Man-Computer Symbiosis,” that humans and computers would develop a close, long-term, and mutually beneficial relationship that would result in more optimal decision-making.6 

Jens Rasmussen

A Danish engineer and psychologist who made significant contributions to the field of safety science, accident research, and human error. To understand at which stage humans should be kept in the loop when interacting with automated systems, he developed the Skill-Rule-Knowledge framework that outlined three levels of cognitive control. Rasmussen advocated that while computers and machines could operate skill- and rule-based tasks, humans should be involved in knowledge-based tasks and intervene if an error occurred.7 

Linus Torvalds

A Finnish software engineer best known for creating Linux, one of the most influential open-source operating systems. While Linux itself is not a human-in-the-loop (HITL) system, the collaborative, open-source model it popularized laid important groundwork for HITL approaches. By enabling a global community of programmers to review, test, and improve code, Linux demonstrated how human feedback could be integrated at scale into the development of complex systems. Torvalds went on to work with a microprocessor manufacturer and later as a project coordinator for the Open Source Development Labs.8 

behavior change 101

Start your behavior change journey at the right place

Impacts

HITL systems have transformed the way AI is applied across multiple fields, combining computational power with human expertise to improve accuracy, reliability, and accessibility. By keeping humans involved in critical decision-making or data validation, these systems ensure that AI tools are both powerful and responsible.

Medical Imaging

In healthcare, AI can be leveraged to analyze vast amounts of data to predict patients at risk of certain conditions, identify cancerous tissue or bone fractures, and develop personalized treatment plans. Not only can this reduce practitioners’ workloads, but often, it can lead to more accurate diagnoses and treatments.

However, although AI tools can sometimes outperform doctors, because errors in healthcare can mean the difference between life and death, it’s important that humans are kept in the loop. AI systems can analyze X-rays to detect tumors, heart conditions, or cancer, but medical professionals should validate these. Doctors will have a better understanding of patient context, a nuance that is required for an accurate diagnosis. By reviewing AI-generated findings, practitioners can intervene and catch misdiagnoses early, providing more information to the algorithm to improve. A study found that while AI tools achieved a 94.6% accuracy in diagnosing breast cancer, when human doctors were incorporated into the process, accuracy improved to 99.5%, demonstrating the benefit of HITL systems.1

Improving accessibility of AI assistants

Many of us now use AI agents, like Siri, Alexa, or Google, to understand and respond to our spoken commands—whether it is asking the weather, adjusting smart home–controlled appliances, or assisting us with hands-free texting and calling. While AI assistants make technology more accessible and convenient, they rely on HITL systems like any other natural language processing tool. To function, they first have to be trained on vast amounts of human-labeled voice and text data to enhance their understanding of language. To handle voice commands in multiple languages and dialects, linguists are often tasked with annotating and validating data, allowing the technology to be more accessible for users across the world.

In 2019, The Common Voice project was developed to open up and decentralize speech technology. As it can be expensive or difficult to get training data for all languages—especially those only used by a minority—the project crowdsourced a multilingual dataset of transcribed speech to train speech technology. Contributors would record their voice, which would then be verified by other users for accuracy. This allowed the researchers to create testing sets for various languages and make AI assistants more accessible to diverse language speakers.9

Improving accuracy of fraud detection

AI has improved fraud detection in banking, as it is capable of analyzing vast amounts of historical data to learn the difference between legitimate and fraudulent transactions, allowing it to identify and block suspicious activities. As scammers are getting smarter and more creative, AI’s ability to notice subtle patterns and anomalies that humans miss helps the tools stay ahead of emerging fraud tactics.10

However, because human financial behavior can be complex, human analysts are kept in the loop. Some AI systems undergo supervised learning, where human-labeled examples of fraudulent behavior are mixed with normal financial records to teach the system what is considered normal versus suspicious. Additionally, analysts will often review AI-flagged cases to determine if they are in fact fraudulent before a customer is alerted, to minimize customer dissatisfaction associated with locked accounts or blocked transactions based on false positives. Customers also often have a chance to provide their feedback and act as a final check when they receive a text message asking them to confirm a transaction.1,10

Controversies

HITL systems bring undeniable benefits, but they also generate trade-offs and ethical questions. By involving humans in AI decision-making, we gain accuracy and oversight—but also slower processes, potential bias, and heightened privacy concerns.

Humans slow down AI

The power of AI comes from its ability to collect and analyze vast amounts of data at a rapid pace, allowing it to detect patterns, identify anomalies, and make predictions that guide decision-making in ways that would be impossible—or extremely time-consuming—for humans to achieve. However, HITL systems are often used for complex situations, where human expertise or knowledge is required. While this can improve accuracy, and also build trust with users who still trust humans more than AI, it also slows down the systems and makes them less scalable. HITL systems are a trade-off between accuracy, latency, cost, and scalability.11 

As keeping humans in the loop slows AI down, it’s important that we identify where human insight is truly needed, and when a fully automated system can handle the task. Rasmussen identified that humans should be kept in the loop for knowledge-based cognitive tasks, while machines should be leveraged for skill- and rule-based cognitive tasks, removing some bottlenecks. 

Is it ethical to keep humans-in-the-loop?

While HITL systems allow AI assistants to understand language better and respond to user queries, this raises the question of whether it is ethical for humans to have access to what may be presumed to be private conversations between a person and their voice-controlled assistants. 

In 2019, privacy concerns were raised when Bloomberg published a report stating that Amazon contracts thousands of people worldwide to transcribe and annotate recordings of clips that users have with Alexa, the company’s digital assistant, in order to improve the tool’s accuracy. Even more troubling is that the report suggested employees sometimes share these snippets internally, either for help identifying the words, or if they hear something troubling. Having humans-in-the-loop means that there may be someone sitting at a computer, listening to a conversation you had with your phone or smart speaker in the confines of your home. While users can delete their voice recordings, or opt out of having their voice requests analyzed, many people do not read the fine print and are unaware that their conversations are being replayed.12 Although the data is anonymized, it may include enough information to identify the user, which raises questions about whether it is ethical to keep humans-in-the-loop in such systems. 

Do HITL systems perpetuate or reduce bias?

The debate as to whether AI perpetuates or reduces bias is ongoing, as is the role of keeping humans in the loop of these systems. Unfortunately, it’s near-impossible for humans to function without bias. As we encounter so much information on a daily basis and have to make rapid decisions without the full context, bias exists because we have to make snap judgments to simplify information processing. As AI can process more information quickly, some people believe that it reduces bias. However, as it is trained on human data, its decisions often continue to perpetuate historical bias. So, when humans are kept in the loop, does it limit or enhance bias?

On the one hand, when humans are involved in supervised learning, they can purposefully train AI on datasets from diverse populations and scenarios that may not be well represented in historical data. For example, many facial recognition datasets historically overrepresented lighter-skinned and male faces, which makes the tools better at recognizing certain races and genders. By correctly labeling darker-skinned and female faces, humans can better train the AI tool to recognize these individuals.13

On the other hand, keeping humans in the loop can also lead to subjective judgments in the labeling of training data, which reintroduces bias. For example, if a human reviews a transcript of a conversation between a customer and a customer assistance bot, to identify if the AI tool correctly identified the emotions of the customer to provide it with the appropriate service response, they may interpret the customers’ tone differently based on their own bias or experience, changing how the data was labeled, therefore feeding the algorithm with biased feedback. 

Case Studies

Mutually beneficial relationship: CAPTCHA

Before accessing a secure website, you may have been asked to participate in a Completely Automated Public Turing test to tell Computers and Humans Apart (known as CAPTCHA) to ensure that you’re not a robot. The goal is to protect websites from spam, abuse, and fraud, by ensuring a human is accessing the site. Users might see an image—or series of images—divided into a grid and be asked to identify which squares contain a vehicle or traffic lights. 

While CAPTCHA’s primary function is to block bots, it is an example of a HITL system. People are being asked to identify objects, thereby labeling data, which can then be used to train AI systems. For example, the labeling of street signs could be used to train autonomous vehicles. Although companies like Google haven’t publicly disclosed that feedback from these tests is being used to train AI, the fact that bots are beginning to be able to solve CAPTCHA tasks shows that they clearly are learning from human input. In fact, research from Microsoft indicates that bots can get CAPTCHAs correct between 85–100% of the time, compared to only 50–85% for humans.14 

Improving AI technology for smart glasses

Over the past decade, various tech companies have released “smart glasses,” wearable computers in the form of eyeglasses. Smart glasses integrate electronic components that allow them to overlay digital information onto the environment someone is looking at, such as providing directions, displaying texts, or allowing you to take pictures of your surroundings. 

But these smart glasses integrate human and technology beyond merging the two physically. Smart glasses’ AI systems rely on humans to label objects. In initial training, humans label objects in datasets to help the technology learn to recognize these in the real world. Once on the market, users’ interactions—such as correcting misidentified objects or confirming AI suggestions—serve as ongoing feedback, allowing the system to learn from real-world scenarios and continuously improve its accuracy and reliability.15

Related TDL Content

Humans and AI: Rivals or Romance?

The relationship between humans and AI has always been contentious. While AI can automate tasks, making us more efficient, advancements in technology also cause fear of job displacement and a lack of control. However, we may not have a reason to fear robots taking over: evidence shows that keeping humans in the loop can lead to the most accurate systems. This article delves into our relationship with AI and makes the case that humans and machines should work together for optimal outcomes.

Why Machines Will Not Replace Us

While there are some tasks that require humans to be kept in the loop—such as nonstandardized tasks—there are other tasks that can be automated to improve efficiency. While automation can cause some skills to become redundant, they also open up the door for us to focus on uniquely human skills. In this article, our writers explore why our fear that machines will put everyone out of work is unwarranted. 

Sources

  1. Macgence. (2025, April 16). Top use cases of human-in-the-loop (HITL) in machine learning and AI. Medium. https://medium.com/@macgenceai/top-use-cases-of-hitl-in-machine-learning-and-ai-a3ca252c9ff5 
  2. Quickcode. (2023, October 2). Why human-in-the-loop? Quickcode. https://quickcode.ai/video/why-human-in-the-loop/
  3. Licklider, J. C. R. (1960). Man–computer symbiosis. IRE Transactions on Human Factors in Electronics, HFE-1(1), 4–11. https://doi.org/10.1109/THFE2.1960.4503259
  4. French Hallway. (2024, January 23). The SRK model: Skill, rule, knowledge. French Hallway. https://frenchhallway.com/blog/2024-01-21-the-srk-model-skill-rule-knowledge/
  5. Amazon Mechanical Turk. (2015, May 21). Bringing future innovation to Amazon Mechanical Turk. Happenings at MTurk. https://blog.mturk.com/bringing-future-innovation-to-mechanical-turk-c67e489e0c37
  6. The Editors of Encyclopaedia Britannica. (2025, August 26). J.C.R. Licklider. Encyclopaedia Britannica. https://www.britannica.com/biography/J-C-R-Licklider
  7. Rasmussen, J. (n.d.) Jens Rasmussen (1926–2018). Retrieved August 26, 2025, from https://www.jensrasmussen.org/
  8. The Editors of Encyclopaedia Britannica. (2025, August 13). Linus Torvalds. Encyclopaedia Britannica. https://www.britannica.com/biography/Linus-Torvalds
  9. Ardila, R., Branson, M., Davis, K., Henretty, M., Kohler, M., Meyer, J., Morais, R., Saunders, L., Tyers, F. M., & Weber, G. (2019). Common Voice: A massively-multilingual speech corpus (Version 2) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.1912.06670
  10. Flinders, M., Smalley, I., & Schneider, J. (2025, April 30). AI fraud detection in banking. IBM Think. https://www.ibm.com/think/topics/ai-fraud-detection-in-banking
  11. Greyling, C. (2025, June 13). Trade-off dynamics of AI agents. Medium. https://cobusgreyling.medium.com/trade-off-dynamics-of-ai-agents-ba4299760d7d
  12. Eadicicco, L. (2019, April 15). Amazon workers reportedly listen to what you tell Alexa — here's how Apple and Google handle what you say to their voice assistants. Business Insider. https://www.businessinsider.com/how-amazon-apple-google-handle-alexa-siri-voice-data-2019-4
  13. Fergus, R. (2024, February 29). Biased technology: The automated discrimination of facial recognition. ACLU of Minnesota. https://www.aclu-mn.org/en/news/biased-technology-automated-discrimination-facial-recognition
  14. Klein, C. (2024, February 18). The truth about those annoying CAPTCHA tests. Scienceline. https://scienceline.org/2024/02/the-truth-about-those-annoying-captcha-tests/
  15. Zulhusni, M. (2025, August 15). Human-in-the-loop work drives AI powering Alibaba’s smart glasses. Artificial Intelligence News.https://www.artificialintelligence-news.com/news/human-in-the-loop-work-drives-ai-powering-alibabas-smart-glasses/

About the Author

Emilie Rose Jones

Emilie Rose Jones

Corporate Communications Manager, TD

Emilie currently works in Marketing & Communications for a non-profit organization based in Toronto, Ontario. She completed her Masters of English Literature at UBC in 2021, where she focused on Indigenous and Canadian Literature. Emilie has a passion for writing and behavioural psychology and is always looking for opportunities to make knowledge more accessible. 

About us

We are the leading applied research & innovation consultancy

Our insights are leveraged by the most ambitious organizations

Image

“

I was blown away with their application and translation of behavioral science into practice. They took a very complex ecosystem and created a series of interventions using an innovative mix of the latest research and creative client co-creation. I was so impressed at the final product they created, which was hugely comprehensive despite the large scope of the client being of the world's most far-reaching and best known consumer brands. I'm excited to see what we can create together in the future.

Heather McKee

BEHAVIORAL SCIENTIST

GLOBAL COFFEEHOUSE CHAIN PROJECT

OUR CLIENT SUCCESS

$0M

Annual Revenue Increase

By launching a behavioral science practice at the core of the organization, we helped one of the largest insurers in North America realize $30M increase in annual revenue.

0%

Increase in Monthly Users

By redesigning North America's first national digital platform for mental health, we achieved a 52% lift in monthly users and an 83% improvement on clinical assessment.

0%

Reduction In Design Time

By designing a new process and getting buy-in from the C-Suite team, we helped one of the largest smartphone manufacturers in the world reduce software design time by 75%.

0%

Reduction in Client Drop-Off

By implementing targeted nudges based on proactive interventions, we reduced drop-off rates for 450,000 clients belonging to USA's oldest debt consolidation organizations by 46%

Read Next

Notes illustration

Eager to learn about how behavioral science can help your organization?