What is AI in Healthcare Equity?
Artificial intelligence (AI) in healthcare equity refers to the ways in which AI systems and algorithms that support medical decisions affect fairness in healthcare. As hospitals and doctors increasingly use AI to diagnose diseases, suggest treatments, and decide how to spend resources, new problems emerge. Sometimes, these AI tools work better for certain patient groups than others, creating unintended bias and differences in treatment. Understanding these issues matters, as artificial intelligence in healthcare is becoming more and more widespread.
The Basic Idea
Maria schedules her annual mammogram at the same clinic she's visited for eight years. When she arrives for her examination, the technician mentions that they now use an AI system to help radiologists detect abnormalities earlier. "It's really advanced," he explains. "Catches things we might miss."
Three days later, Maria receives a callback requesting additional imaging. The AI flagged something suspicious. During her follow-up appointment, the radiologist shows her the scan and explains that the AI identified an area of concern that led to this precautionary step.
What Maria doesn't know is that the computer learned to spot problems by looking at thousands of mammograms from other women. But most of those women had lighter skin and different breast tissue patterns. For women like Maria, whose characteristics differ from the majority of data used to train the algorithm, the computer gets confused more often and sees problems that aren't really there. This means more scary phone calls and extra appointments that she doesn't actually need.
Meanwhile, across town, Robert sits in his doctor’s office as she looks at computer-generated suggestions for his diabetes treatment. The program recommends specific medications based on Robert's test results and medical history, and the doctor has learned to trust these suggestions because they seem thorough and scientific.
However, the algorithm is based on patients who took their medications regularly, never missed appointments, and had stable living situations. Robert sometimes skips his medicine when money is tight, and has missed appointments because he can't take time off work. Because the doctor relies too heavily on an algorithm that doesn’t account for these factors, she is unaware that the suggestions may not be suitable for Robert’s situation.
In both cases, computer systems designed to help people get better healthcare accidentally make things worse. Maria gets scared and has to come back for tests she doesn't need. Robert gets treatment advice that isn’t tailored to his real-life situation.
These stories illustrate how medical programs designed to improve patient care can accidentally make healthcare unequal. The technology itself isn't good or bad. What matters is the information it learned from, where it is applied, and how doctors and patients respond to it—all of which determine whether it benefits everyone equally or makes existing problems worse.
People’s trust in medical technology varies, often because of past experiences with healthcare.1 Some communities have good reason to be cautious, while others welcome new tools. Doctors also react differently to computer suggestions, and those reactions can shape whether the technology helps or harms. The challenge is bigger than fixing computer code. To make patient care more equitable, we need to look closely at each potential impact of AI.
“Bias is a human problem. When we talk about 'bias in AI,' we must remember that computers learn from us."
— Michael Choma, American physician and AI engineer2
Key Terms
Artificial Intelligence (AI): Computer systems designed to perform tasks that typically require human intelligence.
Turing Test: A pioneering experiment by English researcher Alan Turing in his landmark 1950 paper “Computing Machinery and Intelligence.” The test was an early method to assess whether a machine could exhibit human-like intelligence.
Algorithms: A set of step-by-step instructions or rules used by computers to process data and make decisions. In healthcare AI, algorithms can match symptoms to diagnoses, suggest treatments, or prioritize patients—depending on how they’re designed and what data they’re trained on.
Machine Learning: A type of AI that enables systems to learn from data and improve their performance without being explicitly programmed. In healthcare, machine learning models can identify patterns in medical records, predict patient outcomes, or assist in clinical decision-making based on what they’ve learned from past examples.
Ethical AI: The practice of designing and using AI systems in ways that promote fairness, accountability, and transparency while safeguarding individual rights. In healthcare, this includes actively addressing bias, ensuring explainability, and protecting patient autonomy.
History
Artificial Intelligence (AI) first entered the chat in 1950, when British mathematician and computer scientist Alan Turing proposed a provocative question: Can machines think? He introduced a method that became known as the Turing Test—a scenario where a human interacts with both a machine and another human through typed conversation. If the machine’s responses were indistinguishable from the human’s, the machine was said to have passed the test.3
AI systems specifically designed for medicine began to appear in the 1970s, when researchers attempted to encode clinical reasoning into software. These early systems were rule-based, relying on if-then logic written by domain experts—essentially codifying clinical knowledge into structured decision trees. For example:
- MYCIN, developed at Stanford by computer scientist Edward Shortliffe, recommended antibiotic treatments based on user input and explained its rationale to support clinician understanding. It might recommend penicillin if the infection was caused by a specific type of organism and the patient had no known allergy to the drug.4
- INTERNIST-1 and QMR focused on internal medicine diagnoses, linking symptoms to disease profiles using weighted logic. They were deemed by the authors to be insufficient, and not to be relied upon for diagnoses due to various weaknesses.5
- CASNET, created for glaucoma management, modeled disease progression and suggested next steps such as medication or surgery depending on patient data. The framework included three main inputs: disease classifications, patient observations, and pathophysiological states.6
- DXplain offered diagnostic suggestions alongside educational feedback. The inclusion of comments and criticisms from physician users reinforced its recommendations with clinical reasoning.7
These early AI systems laid the foundation for clinical support and guidance in medical decision-making, but were still heavily constrained by the limits of human-authored rules. If a human didn’t provide the explicit rule or critique, the AI wasn’t of much use, and the errors were still far too consistent for the advice to be relied upon. As medical data expanded, newer tools began to shift from manually written rules to algorithms that could learn patterns from real-world inputs.
By the 2000s, the widespread adoption of electronic health records (EHRs) introduced massive datasets, and with them, a new generation of machine learning models. These systems no longer had to rely solely on expert-encoded knowledge because they learned patterns directly from clinical data.
In the 2010s, IBM’s Watson, a question-answering AI system capable of processing natural language, shifted into healthcare after its high-profile Jeopardy! win. It was tasked with literature searches and clinical decision support, and IBM purchased massive amounts of clinical data as part of its foray into the field.8 Meanwhile, the COVID-19 pandemic brought urgency to virtual triage, remote monitoring, and AI-powered chatbots. These tools brought promise and optimism in a dark time because they could scale quickly under pressure, but were still experiencing challenges as they attempted to adapt to new local characteristics—each with their own unique demographics, languages, equipment, and protocols.9
Today, AI tools are increasingly integrated into healthcare practices. Although researchers and computer engineers are continuously working to improve their reliability, these challenges persist.
People
Alan Turing
A highly influential figure in the field of AI, beginning with his imitation game—now called the Turing test—in 1950. Turing was an English researcher prominent in several fields, including math, philosophy, and cryptanalysis, who is considered the father of theoretical computer science.
Edward Shortliffe
An American computer scientist and physician who, in the 1970s, created MYCIN—one of the first rule-based AI systems used in medicine. Shortliffe is considered the pioneer of biomedical informatics: the field of study dedicated to the improvement of healthcare through the integration of information and computer systems.
Watson
This AI system, created by IBM, became famous in 2011 after winning as a contestant on the American game show Jeopardy! Once seen as the future of AI in medicine, IBM’s Watson for Oncology failed to meet expectations in clinical practice. It made unsafe treatment recommendations and couldn’t adapt to regional care standards or evolving data, exposing the risks of deploying high-profile AI tools without clear alignment to healthcare realities.
behavior change 101
Start your behavior change journey at the right place
Impacts
AI tools are shaping healthcare access, quality, and outcomes—but without a focus on equity, these systems risk reinforcing existing disparities. Whether through biased data, opaque algorithms, or unequal implementation, AI can widen the very gaps it promises to close. Understanding its impacts on equity is critical for developers, clinicians, healthcare leaders, and government policymakers who aim to ensure these technologies serve all patients fairly.
Bias and health equity
AI systems in healthcare can unintentionally disadvantage patients from marginalized groups when the data they are trained on fails to represent everyone equally. This can lead to uneven diagnostic accuracy, misallocation of resources, and unequal healthcare outcomes across racial, ethnic, gender, or socioeconomic lines.
Research has documented multiple cases where bias manifests in AI tools. A landmark 2019 study found that a widely-used algorithm underestimated Black patients’ health needs by using healthcare spending as a proxy, leading to systematically fewer referrals despite similar needs.10 In dermatology, AI performance dropped substantially on images of darker skin tones compared to lighter ones, indicating significant disparities in diagnostic accuracy.11 Similarly, mental health assessment models are more than three times less predictive for Black users than for White users when analyzing social media language patterns to assess depression risk.12 These examples underscore how skewed data can translate into real-world inequities, turning helpful tools into sources of unintended harm.
Trust and transparency in clinical practice
AI systems in healthcare rely on trust to be effective. Patients and clinicians need to know how these tools work, how decisions are made, and where the data comes from. If the process isn’t clear or the data lacks representation, trust weakens and use declines—especially among populations that have historically been underserved, misdiagnosed, or mistreated by the healthcare system.
Transparent design and clear communication can strengthen that trust. A 2022 research study found that explaining data sources, limitations, and intended use helps healthcare professionals make informed choices.13 This kind of openness gives clinical teams a stronger foundation for deciding when and how to use AI systems appropriately and ethically.
Trust also grows through day-to-day practice. In research involving cancer care, clinicians trusted AI more when there were shared routines, team-based checks, and clear communication.14 When developers and clinical leaders prioritize open systems and collaborative environments, it becomes easier to build confidence in AI over time.
Accountability and liability
When AI tools are involved in diagnosis, treatment, or decision support, questions arise about who is responsible when something goes wrong. In systems that are constantly learning or that operate as “black boxes,” it becomes harder to identify the source of an error—whether it stems from a flawed algorithm, poor data, or human misuse.
This uncertainty creates risk for clinicians, health systems, and patients. Without clear accountability structures, it’s harder for individuals to challenge unfair decisions or seek redress. Clinical teams may also hesitate to use AI tools if liability remains vague or poorly defined. Research shows clinicians who normally accept and share accountability through explanations of their assessments are hesitant to adopt AI and have reduced confidence in its safety, especially when the AI fails to sufficiently support its stance or diagnoses.15
Clear standards for documentation, oversight, and error reporting help reduce liability concerns and increase transparency. When systems are designed with built-in audit trails and explained functionality, it becomes easier to determine what went wrong and who should respond.
Controversies
AI has the potential to improve healthcare access and outcomes—but without close attention to its design and deployment, it can also deepen existing inequities. These controversies reflect the unresolved tensions shaping how, where, and for whom AI serves in healthcare.
Bias mitigation vs. bias amplification
AI systems in healthcare are often built with the intention of correcting human error or overcoming biased clinical decisions. However, many models trained on historical data inadvertently replicate or amplify the same systemic inequalities they were meant to avoid.
Bias in healthcare predates AI. A growing body of evidence shows that even widely used medical devices can perpetuate inequity when they are not designed or validated for all populations. For example, a 2021 study found that pulse oximeters—used to measure blood oxygen levels—were three times more likely to miss dangerously low oxygen levels (occult hypoxemia) in Black patients compared to White patients, even when readings appeared normal.16 This underdetection can delay interventions for conditions like COVID-19 or pneumonia, directly impacting patient safety and outcomes.
This case is just one of many where biased proxies and assumptions result in skewed outputs. The controversy centers on whether AI can ever be "neutral" when it learns from a healthcare system built on structural racism, underrepresentation, and unequal access. While researchers have suggested third-party or government audits to identify and correct for some biases, critics argue that many of these fixes are superficial, and that the underlying logic of prediction—optimizing for efficiency, cost savings, or population-level outcomes—often conflicts with individual and community needs.17
Calls for equity-aware machine learning frameworks continue to grow, but implementation remains uneven. And without transparency in how fairness is defined or measured, the risk persists that AI will entrench bias rather than reduce it—especially for populations who are already medically underserved.
Liability and accountability for AI-driven decisions
One of the most pressing and unresolved controversies in AI and healthcare is who is accountable when algorithmic decisions lead to harm. As AI systems become embedded in clinical workflows—especially in risk assessment, diagnostics, and mental health triage—the legal and ethical frameworks around responsibility remain underdeveloped.
AI-driven tools can influence medical decisions, but they often function as “black boxes” with limited transparency or explainability—that is, clinicians can see the outputs but not fully trace or explain how the system arrived at them.18 If an AI system makes an error—such as underestimating a patient's suicide risk or delaying a referral for psychiatric support—who is legally liable? The clinician? The hospital? The software vendor?
Accountability is even more complicated when vendors treat their algorithms as proprietary. Without transparency or external review, healthcare systems may adopt tools with hidden flaws, and patients may never know how or why they were flagged, misclassified, or overlooked.
Regulatory gaps and uncertainty
AI in healthcare is advancing so quickly that regulatory frameworks built for traditional, static medical devices are struggling to keep pace. Current oversight mechanisms were never designed for adaptive algorithms that can evolve as they process new data, making it challenging to ensure consistent performance, patient safety, and equitable outcomes over time. This dynamic nature means that an AI tool approved today could function differently months or years later without triggering additional review.
The result is a regulatory mismatch that leaves gray areas in accountability, especially around bias detection, transparency, and post-deployment monitoring. Some critics suggest regulators simply must find ways to keep up: provide a framework for defining regulatory scope, ensuring consistent performance across populations, and balancing innovation with patient safety.19 Such governance would need to balance fostering innovation with safeguarding public health. If hospitals are going to lean in on using these tools, regulatory agencies have no choice but to address the necessary governance to prevent potential risks.
Case Studies
Sarah Meredith’s urgent liver transplant
The average wait time in the UK for a liver transplant is 68 days, according to the National Health Service (NHS). Sarah Meredith, a 31-year-old living with cystic fibrosis and other genetic conditions, was forced by an inequitable AI system to wait 25 months. It turns out her access to a liver transplant was being determined by an opaque algorithm known as the Transplant Benefit Score (TBS).20 She learned that the algorithm ranked patients based on a projected five-year survival benefit rather than overall need. This meant that older patients often scored higher under the algorithm because the model weighted short-term survival benefits more heavily than longer-term outcomes. As a consequence, it was nearly impossible for patients under age 45—no matter how serious their condition—to secure prioritization over someone older.
Despite Meredith's attempts to appeal her low score, she encountered resistance from clinicians and NHS officials, with some suggesting she “wouldn’t understand” how the algorithm worked. It wasn’t until she used a public web tool—developed by a researcher studying the TBS—that she could see how algorithmic logic systematically disadvantaged younger patients. Experts confirmed that the system overlooked key aspects such as longer-term benefit and quality of life, reinforcing disparities in access to life-saving care. Even a year after Meredith’s eventual transplant and after increasing public awareness of the critical flaws in the AI algorithm, the battle to remove the system from usage continued.21
U.S. Medicare Advantage AI exploitation (nH Predict)
In the U.S., many health plans use nH Predict—an AI-driven tool from NaviHealth—to estimate patients' post‑acute care needs and inform decisions on approving services. For example, a patient in rehabilitation needing a few extra days of treatment due to a slower pace of recovery would be cut off from payment after the prescribed days authorized by the algorithm. However, a class-action lawsuit filed in November 2023 alleges the system has been used not just to streamline reviews, but to automatically deny care for a significant percentage of patients—particularly those in elderly or vulnerable populations—by making coverage determinations without human oversight.22 If a clinician’s advice would extend a patient’s approved care period, they risked being disciplined, as if missing the target were a performance indicator.
In response to growing concerns over AI-powered coverage denials by some Medicare Advantage insurers, the Centers for Medicare and Medicaid Services (CMS) issued clarifying guidance in 2024. The agency made it clear that while AI and algorithms can assist in coverage determinations, final decisions must be based on each patient’s individual circumstances, in accordance with federal rules (42 CFR § 422.101(c)).23 This ensures clinicians remain responsible for medical necessity decisions and that AI does not replace—nor punish—human judgment.
Related TDL Content
AI alignment
For artificial intelligence to be useful and cause the least harm, it should align with human values and ethics. AI alignment is the design of systems whose goals stay anchored to human (and ethical healthcare) values.
The potential and pitfalls of AI in healthcare
Whether it's relieving clinician burnout or detecting disease earlier, AI holds powerful promise in healthcare. However, pitfalls such as biased data and misaligned incentives can exacerbate disparities if ethics and equity aren’t considered from the start.
Sources
- Hudson, P., & Williams, M. A. (2023, January 31). People are much less likely to trust the medical system if they are from an ethnic minority, have disabilities, or identify as LGBTQ+, according to a first-of-its-kind study by Sanofi. Fortune. https://fortune.com/2023/01/31/people-trust-health-medical-system-ethnic-minority-disabilities-identify-lgbtq-study-sanofi-hudson-williams/
- Miliard, M. (n.d.). Yale study shows how AI bias worsens healthcare disparities. Healthcare IT News. https://www.healthcareitnews.com/news/yale-study-shows-how-ai-bias-worsens-healthcare-disparities
- Turing, A. M. (1950). I.—Computing machinery and intelligence. Mind, 59(236), 433–460. https://doi.org/10.1093/mind/LIX.236.433
- Buchanan, B. G., & Shortliffe, E. H. (1984). Rule-based expert systems: The MYCIN experiments of the Stanford Heuristic Programming Project. Addison-Wesley.
- Miller, R. A., Pople, H. E., & Myers, J. D. (1986). INTERNIST-1, an experimental computer-based diagnostic consultant for general internal medicine. New England Journal of Medicine, 307(8), 468–476.
- Weiss, S. M., Kulikowski, C. A., Amarel, S., & Safir, A. (1977). A model-based method for computer-aided medical decision-making. Artificial Intelligence, 8(1), 145–170.
- Barnett, G. O., Cimino, J. J., Hupp, J. A., & Hoffer, E. P. (1987). DXplain. An evolving diagnostic decision-support system. JAMA, 258(1), 67–74. https://doi.org/10.1001/jama.258.1.67
- O’Leary, L. (2022, January 31). How IBM’s Watson went from the future of health care to sold off for parts. Slate Magazine. https://slate.com/technology/2022/01/ibm-watson-health-failure-artificial-intelligence.html
- Hu, Y., Jacob, J., Parker, G. J., Hawkes, D. J., Hurst, J. R., & Stoyanov, D. (2020). The challenges of deploying artificial intelligence models in a rapidly evolving pandemic. Nature Machine Intelligence, 2(6), 298–300. https://doi.org/10.1038/s42256-020-0185-2
- Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science (New York, N.Y.), 366(6464), 447–453. https://doi.org/10.1126/science.aax2342
- Daneshjou, R., Vodrahalli, K., Novoa, R. A., Jenkins, M., Liang, W., Rotemberg, V., Ko, J., Swetter, S. M., Bailey, E. E., Gevaert, O., Mukherjee, P., Phung, M., Yekrang, K., Fong, B., Sahasrabudhe, R., Allerup, J. A. C., Okata-Karigane, U., Zou, J., & Chiou, A. S. (2022). Disparities in dermatology AI performance on a diverse, curated clinical image set. Science advances, 8(32), eabq6147. https://doi.org/10.1126/sciadv.abq6147
- Rai, S., Stade, E. C., Giorgi, S., Francisco, A., Ungar, L. H., Curtis, B., & Guntuku, S. C. (2024). Key language markers of depression on social media depend on race. Proceedings of the National Academy of Sciences, 121(14). https://doi.org/10.1073/pnas.2319837121
- Benda, N. C., Lovett Novak, L., Reale, C., & Ancker, J. S. (2022). Trust in AI: Why we should be designing for appropriate reliance. Journal of the American Medical Informatics Association, 29(1), 207–212. https://doi.org/10.1093/jamia/ocab238
- Procter, R., Tolmie, P., & Rouncefield, M. (2022). Holding AI to account: Challenges for the delivery of trustworthy AI in healthcare. ACM Transactions on Computer-Human Interaction 30(2), 1-34. https://doi.org/10.1145/357700
- Norori, N., Hu, Q., Aellen, F. M., Faraci, F. D., & Tzovara, A. (2021). Addressing bias in big data and AI for health care: A call for open science. Patterns (New York, N.Y.), 2(10), 100347. https://doi.org/10.1016/j.patter.2021.100347
- Gerke, S., Minssen, T., & Cohen, I. G. (2020). Ethical and legal challenges of artificial intelligence-driven healthcare. Artificial Intelligence in Healthcare, 295–336. https://doi.org/10.1016/B978-0-12-818438-7.00012-5
- Chen, I. Y., Szolovits, P., & Ghassemi, M. (2021). Can AI help reduce disparities in general medical and mental health care? AMA Journal of Ethics, 23(2), E180–E186. https://doi.org/10.1001/amajethics.2019.167
- Cabitza, F., Rasoini, R., & Gensini, G. F. (2017). Unintended consequences of machine learning in medicine. JAMA, 318(6), 517–518. https://doi.org/10.1001/jama.2017.7797
- McKee, M. (2022). The Challenges of Regulating Artificial Intelligence in the Healthcare Sector. Journal of Clinical Decision Support, 11, Article 3159. https://doi.org/10.34172/ijhpm.2022.7261
- Financial Times. (2023, November 9). Algorithms are deciding who gets organ transplants. Are their decisions fair? Financial Times. (Original reporting by Sarah Meredith case).
- Narayanan, A., & Kapoor, S. (2024, November 11). Does the UK’s liver transplant matching algorithm systematically exclude younger patients? AI Snake Oil. https://www.aisnakeoil.com/p/does-the-uks-liver-transplant-matching
- Mello, M. M., & Rose, S. (2024, March 7). Denial—Artificial Intelligence Tools and Health Insurance Coverage Decisions. JAMA Forum. https://doi.org/doi:10.1001/jamahealthforum.2024.0622
- Frequently Asked Questions related to Coverage Criteria and Utilization Management Requirements in CMS Final Rule (CMS-4201-F). (n.d.). https://www.aha.org/system/files/media/file/2024/02/faqs-related-to-coverage-criteria-and-utilization-management-requirements-in-cms-final-rule-cms-4201-f.pdf



















