Revolutionizing Fair Hiring: Navigating Gender Bias in AI Algorithms

Published · By Celine Huang

An evolving hiring landscape

With 92% of HR leaders using or planning to use AI for candidate screening, according to a 2021 Mercer report, the question isn't whether AI will transform hiring—it's how we can ensure it does so fairly. While AI promises efficiency, its potential to amplify existing biases, especially against women in male-dominated fields, poses a critical challenge for employers across all industries. This problem isn’t new, either; in 2018, Amazon faced significant public scrutiny when gender bias was discovered in its resume screening algorithms.1

Luckily, all hope is not lost. With awareness of algorithmic bias growing, researchers are refining tools to measure and counteract unintentionally discriminatory outcomes. But how do different debiasing strategies shape not only fairer hiring decisions, but also whether candidates choose to apply at all? It's a question that gets to the core of workplace diversity and inclusion, moving beyond simple metrics to examine how AI shapes the very makeup of our applicant pools. To answer this question, Dr. Edwin Ip from the University of Exeter tested how awareness of gender bias and various debiasing methods changed application decisions in a simulated hiring experiment with over 700 U.S. adults.

Methods

Defining fairness

Before diving into the experiment itself, let’s explore what fairness in hiring actually means. While no algorithm is perfect, an algorithm is considered biased if it makes repeated, systematic errors about particular groups. Dr. Ip tested three common debiasing approaches, each addressing different aspects of fairness with unique trade-offs between statistical accuracy, group representation, and perceived fairness:

  1. Equality of accuracy: ensures that qualified candidates have equal chances of selection regardless of gender, calibrating the algorithm's precision across demographic groups. 
  2. Equality of predictions: maintains equal selection rates between men and women, directly addressing representation in the candidate pool. Women are the least disadvantaged in this approach, as it increases female representation among successful applicants.
  3. Gender blindness: completely removes gender-related information from the decision-making process

Examining effects on application behavior

To systematically evaluate AI fairness in hiring, 736 working adults from across the United States participated in an online experiment. These participants—a balanced mix of men and women with diverse educational backgrounds—first provided details about their education, work experience, top skills, and self-assessed quantitative ability, which were compiled into résumé-style profiles for algorithmic evaluation. Then, they were asked to “apply” for one of two job options: Job A, a competitive and high-paying position where 20% of applicants would be selected by a hiring algorithm, and Job B, a non-competitive and lower-paying position where 50% of applicants would be selected at random. 

Selection for Job A was evaluated under four different algorithms: the three aforementioned debiased methods and a nondebiased “best fit” model that maximized accuracy. Participants were given information about potential biases in the hiring algorithms, then asked to choose between Job A and B, encountering each possible hiring algorithm once in a random order. A subgroup of participants was also asked to choose between Job A and Job B before being given information about possible gender bias, so the researchers could examine the effects that this information had on their choices. 

Afterwards, participants completed an assessment on quantitative skills to determine the most qualified applicants of the bunch, information that was used by Dr. Ip for analysis but was not made known to the hiring algorithms. Participants also rated how fair they perceived each algorithm, their preference for each evaluation method, and their beliefs about their chances, to gain insight into the psychological mechanisms behind their choices.

Key Findings

When participants learned about potential gender bias in hiring algorithms, women became 18.1 percentage points less likely to apply for competitive positions, while men became 10.9 percentage points more likely. In practice, this means that applicant pools may be skewed by gender even before the algorithms screen candidates, based on an awareness of potential bias alone. Further, all three debiasing algorithms significantly increased female participation without changing the overall quality of the applicant pool. 

Between each approach, the equality of prediction approach was most effective at attracting female applicants, while the gender-blind model earned high marks for perceived fairness from both men and women. On the other hand, men were most likely to apply for Job A when it used the nondebiased approach, the condition where they had the most advantage. 

behavior change 101

Start your behavior change journey at the right place

Implications

With these insights, companies across sectors (particularly in tech and finance, where gender gaps persist) can intentionally design fairer hiring processes without lowering applicant quality. As Dr. Ip outlines, organizations can choose between different approaches based on their goals. If the focus is solely on diversity, employers can use an algorithm that selects an equal proportion of male and female applicants. Conversely, gender blind algorithms can build trust and foster inclusion, even if they don’t necessarily maximize the proportion of female applicants.  

Still, it's important to acknowledge the study's limitations. While it focused specifically on gender bias, future research needs to examine how these algorithms affect other forms of diversity, such as race and background. Additional studies must also examine how various factors intersect and interact with each other, such as race and gender. In sum, there's still much to learn about creating truly inclusive AI systems.

What now?

When it comes to hiring algorithms, fairness isn't just an ethical imperative—it's a powerful tool for attracting and identifying the best talent, regardless of gender. In a world where AI increasingly influences who gets hired, this study demonstrates how technology can either reinforce old biases or help break them down. By choosing the right approach to debiasing, organizations can create hiring processes that don't just talk about equality—they deliver it.

This article summarizes:

Ip, E. (2025). Fair AI in hiring: Experimental evidence on how biased hiring algorithms and different debiasing methods affect the quality and diversity of applicants. Behavioral Science & Policy, 11(1), 44-54. https://doi.org/10.1177/23794607251353585 (Original work published 2025)

Additional References

Goodman, R. (2023, February 27). Why Amazon’s Automated Hiring Tool discriminated against women: ACLU. American Civil Liberties Union. https://www.aclu.org/news/womens-rights/why-amazons-automated-hiring-tool-discriminated-against

About the Author

Celine Huang

Content Lead

Celine Huang is a Summer Content Intern at The Decision Lab. She is passionate about science communication, information equity, and interdisciplinary approaches to understanding decision-making. Celine is a recent graduate of McGill University, holding a Bachelor of Arts and Sciences in Cognitive Science and Communications. Her undergraduate research examined the neurobiology of pediatric ADHD to improve access to ADHD diagnoses and treatments. She also sits on the North American Coordinating Committee of Universities Allied for Essential Medicines (UAEM), where she applies her behavioral science background to health equity advocacy. In her free time, Celine is an avid crocheter and concertgoer.

About us

We are the leading applied research & innovation consultancy

Our insights are leveraged by the most ambitious organizations

Image

“

I was blown away with their application and translation of behavioral science into practice. They took a very complex ecosystem and created a series of interventions using an innovative mix of the latest research and creative client co-creation. I was so impressed at the final product they created, which was hugely comprehensive despite the large scope of the client being of the world's most far-reaching and best known consumer brands. I'm excited to see what we can create together in the future.

Heather McKee

BEHAVIORAL SCIENTIST

GLOBAL COFFEEHOUSE CHAIN PROJECT

OUR CLIENT SUCCESS

$0M

Annual Revenue Increase

By launching a behavioral science practice at the core of the organization, we helped one of the largest insurers in North America realize $30M increase in annual revenue.

0%

Increase in Monthly Users

By redesigning North America's first national digital platform for mental health, we achieved a 52% lift in monthly users and an 83% improvement on clinical assessment.

0%

Reduction In Design Time

By designing a new process and getting buy-in from the C-Suite team, we helped one of the largest smartphone manufacturers in the world reduce software design time by 75%.

0%

Reduction in Client Drop-Off

By implementing targeted nudges based on proactive interventions, we reduced drop-off rates for 450,000 clients belonging to USA's oldest debt consolidation organizations by 46%

Notes illustration

Eager to learn about how behavioral science can help your organization?