The Big Problem
You’ve spent months building it, a mobile therapy app designed to support people when they need it most. The features are polished. The content is trauma-informed. The user flow is clean. Before launch, you run a pilot focus group to test the experience—gather feedback, catch blind spots, and adjust where needed. But something still nags at you. Is the feedback in your pilot group honest, or just agreeable? Who’s missing from the room entirely? And what about the people who never download mental health apps at all—would they possibly consider downloading this app? Groupthink, selection bias, and drop-off rates are familiar problems with significant consequences. Whether you're testing a digital health tool, launching a climate-conscious product, or piloting a public policy campaign, the challenge stays the same: how do you gather real insight before it's too late to change course?
Synthetic focus groups are stepping into that gap. With the help of large language models (LLMs), researchers can simulate diverse perspectives, pressure-test ideas, and explore blind spots that traditional methods tend to miss. Done well, they don’t replace humans, they help researchers hear more of them. Behavioral science helps make synthetic focus groups usable, responsible, and ready for the real world. Especially when synthetic tools promise faster timelines and broader reach, those principles help ensure we’re not just hearing more voices. They help ensure we’re hearing what matters.
TL;DR
- Traditional and virtual focus groups still face bias, selection issues, and logistical barriers. Artificial Intelligence (AI) might ease these limits, yet its potential for true human collaboration remains largely untapped.
- Hybrid human–AI focus groups can create a new model for inclusion, combining AI’s pattern-recognition power with the lived experience of real communities. This collaboration turns participation into an iterative exchange, with humans correcting what machines surface.
- Behavioral auditing keeps synthetic focus groups accountable by embedding human critique into AI workflows, preventing overconfidence, questioning bias, and ensuring it gets corrected and documented in real time.
What are Synthetic Focus Groups?
In this article, we’ll be focusing on synthetic focus groups—AI-generated simulations of group discussions designed to complement, not replace, human research. These models draw on large text and speech datasets to approximate how diverse participants might respond to a given topic. They’re meant to spot patterns, test ideas, and surface blind spots before the real conversations begin. They may be fast and efficient, but without human interpretation, their findings can drift from the realities they aim to reflect.
Why Traditional and Virtual Focus Groups May Benefit from AI
Focus groups have always been a workhorse for qualitative insight but their success might depend less on the protocol used and more on the people in the room, both answering questions and facilitating discussion. Discussion is an assisted process: one voice moderates, several respond, and the whole exercise hinges on balance. When a facilitator keeps conversation flowing, insights surface naturally. When they lose control, a single outspoken participant can dominate the floor, skewing the discussion and muting quieter perspectives.1,2 It’s a reminder that talk is data, but only when everyone’s actually talking.
Recruitment can also compound the issue. Most participants are self-selected, introducing selection bias, a distortion that occurs when volunteers differ systematically from those who stay out of the research process.3 Convenience replaces representation; flexibility replaces diversity. Those juggling multiple jobs or caregiving responsibilities are less likely to devote their time to a mid-week session.4,5 The data then reflect who could attend, not who should be heard. What looks inclusive on a spreadsheet may, in practice, reinforce the same structural blind spots decision-makers hope to correct.
Virtual formats were expected to help solve these problems with in-person designs. The shift after COVID-19 expanded reach, enabling participation across time zones and reducing social pressure through anonymity.6,7 Still, digital sessions have their own friction. Time differences make scheduling messy; lag interrupts conversation flow; and asynchronous text-based groups often trade richness for speed. Responses come in shorter bursts, stripped of tone or nuance, and the dialogue loses its rhythm.8,9 Convenience, it turns out, doesn’t always translate into insight.
LLMs might help with these logistical barriers. Systems like ChatGPT can simulate varied viewpoints, generate counterarguments, and pilot early hypotheses before live recruitment begins.7,10,11 They won’t replace human dialogue, nor should they, but they could lighten the load, especially when they’re used alongside people throughout the research process, from study design to data interpretation. With thoughtful integration, synthetic and human insight may start working in concert rather than competition.
Challenge #1: The Hidden Costs of Unequal Representation in Research
Across industries, evidence drives decisions, but the evidence itself isn’t always built on everyone’s experience. Whether we’re designing healthcare programs, digital products, or public policy, the people most affected by our choices often aren’t in the data that guides them. They’re not unwilling; they’re just not being reached. Synthetic focus groups promise to change that—to represent the unrepresented, to give voice to perspectives that real recruitment can’t reach. However, their foundation is human data, and that foundation is already uneven.
In healthcare, the gap’s impossible to ignore. Underserved populations—ethnic minorities, low-income families, rural communities, people with disabilities—carry the heaviest burden of illness yet appear the least in the studies meant to improve their outcomes.12,13 The same thing happens elsewhere: the consumers who rely on public systems or lower-cost services rarely show up in the focus groups that define them.14 We keep producing well-intentioned solutions tailored to the people who were easiest to find. Most researchers already know this—it’s not a secret. Recruiting outside familiar networks takes time, energy, and trust. Hospitals, clinics, and universities are convenient pipelines, but they can also end up becoming echo chambers. Folks who live far from those centers, who don’t read research bulletins or check university websites, might never even hear there’s a study happening. The issue isn’t lack of interest—it’s design.
Familiarity bias can also negatively impact recruitment, keeping researchers close to the networks that already work. These aren’t moral failings—they’re human shortcuts within systems that reward efficiency over reach. The cost, though, is real. When key voices go missing, we don’t just lose diversity, we lose accuracy too. Treatments that look promising on paper might stumble in the real world. Campaigns meant to inform fall flat because the message doesn’t land among the target audience. And across sectors, from clinical care to AI governance, decisions made without the full range of lived experience risk hardening the same inequities they aim to fix.
The real challenge isn’t who we recruit, it’s how we simulate them. Representation can’t be reverse-engineered from biased data; it has to be designed into the system from the start. Behavioral science helps here. Health research might show it most clearly, but the lesson’s universal: if we want smarter evidence—and fairer outcomes—we’ve got to make sure the story we’re telling actually includes everyone in it.
behavior change 101
Start your behavior change journey at the right place
Opportunity #1: Countering Cognitive Biases with Human AI Focus Group Models
Traditional focus groups have always been powerful tools for understanding human perspectives, but they face persistent limits: recruitment bottlenecks, social desirability bias, and underrepresentation of marginalized populations.1,5,14 For decision-makers who rely on qualitative insights, this creates a recurring blind spot—the same voices get heard while others stay out of reach. Hybrid human–AI focus groups can change that by pairing the pattern-finding strength of artificial intelligence with the contextual judgment of real communities.
How Hybrid Focus Groups Work
In a hybrid design, researchers use AI to simulate or “pre-model” focus-group discussions based on publicly available, ethically sourced data such as interviews, community forums, or local media transcripts. The goal isn’t to replace participants but to prototype hypotheses—to see what themes emerge before recruiting real people. Once simulated results are produced, human participants are brought in to validate, challenge, or reinterpret them.
Consider a public-health team studying vaccine hesitancy in rural communities. Instead of starting with empty recruitment lists, researchers train a language model on local radio discussions, community Facebook groups, and news comments. The model surfaces preliminary themes—distrust of centralized health authorities, logistical barriers to access, or confusion about eligibility. These become starting points, not conclusions.
Next, researchers present these AI-generated findings to actual community members and ask: “Does this sound accurate? What would you change?” Participants might agree with some points but reject others: “It’s not about distrust—it’s about timing. We get the information too late.” That correction adds context no algorithm could infer. Researchers can then refine their recruitment strategies and messaging based on what truly drives behavior.
The Behavioral Mechanism
This two-stage process helps counteract two cognitive biases that often distort qualitative research:
- Anchoring bias, where researchers overvalue early insights or initial responses.
- Confirmation bias, where analysts interpret new data to fit existing beliefs.
When human participants are explicitly invited to contradict simulated outputs, those biases are disrupted. AI provides a preliminary “anchor,” but the human stage deliberately breaks it. The result is a richer, more accurate model of how communities think, speak, and decide.
Why This Matters
AI contributes what humans can’t:
- Pattern recognition at a non-human scale: mapping co-occurring themes across thousands of data points.
- Speed and traceability: generating transparent records of how insights were derived.
- Early detection of gaps: revealing populations or perspectives absent from existing datasets.
Humans, meanwhile, bring what AI lacks: emotional nuance, ethical discernment, and lived credibility. The hybrid model turns inclusion into an iterative system—AI for discovery, humans for grounding, AI for refinement.
Where It Works
Evidence from other fields already shows the value of pairing machine intelligence with human oversight. In security systems, Cao et al. demonstrated that a human-in-the-loop framework for X-ray luggage inspection improved accuracy by routing uncertain cases to inspectors and feeding their feedback into continuous model retraining.15 Similarly, Humans in the Loop, an award-winning social enterprise, uses active learning to flag ambiguous AI predictions for human audit, reducing bias and model drift.16 These systems work because algorithms handle scale and detection, while people bring context, nuance, and judgment. Translating that principle to qualitative research means treating communities as live auditors of synthetic insights—keeping inclusion honest, iterative, and grounded, because good data’s got to be tested, not trusted.
Challenge #2: Synthetic Models Can’t Recreate Voices Absent from the Dataset
Synthetic focus groups were supposed to fix what traditional research keeps getting wrong—too few voices, too much bias, not enough reach. Feed an algorithm a mountain of transcripts, social posts, and surveys, and it can conjure a roomful of simulated participants in minutes. On paper, it’s diversity on demand. In reality, it’s déjà vu. The same blind spots that haunt real-world research creep back in.
The problem begins where it always does: the data. Most public-health and social-behavior datasets weren’t built to represent everyone; they were built to be convenient. Hospital records capture those who can get to hospitals. Digital health surveys reach the literate, the connected, and the interested. Rural voices, Indigenous communities, and those without reliable access to care often don’t make the dataset at all. When synthetic focus groups learn from that information, they don’t fill gaps—they reproduce them. The result is a simulated conversation that feels inclusive but merely echoes the same narrow demographic chorus.
This is representation bias in action: when certain populations are under-sampled or excluded, and their absence becomes statistically invisible.17 It’s how models end up mistaking “urban” as the default human condition, while rural or marginalized groups are treated as outliers. Then comes measurement bias—the shortcut of using the wrong yardsticks for human reality.18,19 Health researchers, for instance, may rely on hospital attendance as a proxy for access to care or on smartphone use as a proxy for digital literacy. Those indicators mean radically different things in different socioeconomic or cultural contexts. Finally, there’s aggregation bias, which arises when patterns that hold for a population are assumed to apply equally across its subgroups. It’s what happens when differences inside a dataset are hidden by summary statistics. For example, an AI model trained on national vaccination data might find high average uptake overall, yet mask the fact that rates are far lower among those living in rural areas or among specific ethnic groups. When that same data is used to inform synthetic focus groups, it shapes which voices are simulated and which are left faint in the mix—creating the appearance of balanced representation even when some perspectives barely register.
Empirical evidence already shows how these mechanisms interact. A widely used U.S. healthcare-risk algorithm underestimated the needs of Black patients by equating health expenditure with health status—historic bias disguised as precision.20 Sepsis prediction models trained in high-income hospitals underperformed for Hispanic patients because representation in the source data was minimal.21 These aren’t isolated cases; they reveal what happens when models optimize for efficiency rather than equity.
Now extend that logic to qualitative inference. Suppose a government team uses synthetic focus groups to gauge attitudes toward a carbon tax. The model, trained largely on online discourse and national news, generates articulate, policy-savvy participants debating economic competitiveness and consumer prices. The transcript looks balanced, even nuanced. Yet, when researchers later hold real focus groups in manufacturing towns, they find a different concern altogether—fear of job loss, uneven enforcement, rising fuel costs. The simulation didn’t fabricate these voices; it erased them through generalization.
That interpretive gap isn’t purely technical—it’s cognitive. Research on automation bias shows that humans often over-trust computer-generated output because it appears objective or immune to human error.22 Analysts may treat synthetic transcripts as neutral evidence rather than probabilistic inference. Compounding this is anchoring bias: early machine-generated summaries can frame later interpretation, making subsequent human input seem anecdotal by comparison.23 Once these biases take hold, data feels factual, and questioning it feels unnecessary.
The challenge, then, isn’t that synthetic focus groups malfunction—it’s that they can appear too convincing. Their conversational realism gives a false sense of epistemological completeness. Without systematic cross-validation against lived data, the apparent neutrality of AI-generated dialogue can conceal layers of representational distortion. Over time, those distortions may drift into decision frameworks, shaping funding priorities, communication strategies, and even regulatory language.
In short, the risk isn’t technological overreach but interpretive complacency. Synthetic focus groups can replicate the cadence of human discussion, but they can’t reproduce the heterogeneity of human experience. Unless continuously benchmarked against real-world input, they risk turning complex public realities into statistical ventriloquism—voices that sound authentic yet echo the same partial histories we’ve always heard.
Opportunity #2: Embedding Behavioral Auditing in Synthetic Focus Groups
If the problem with synthetic focus groups is overconfidence, the solution isn’t to abandon them—it’s to keep them accountable. Behavioral auditing embeds human challenge directly into the simulation loop. Instead of treating synthetic participants as finished data, it treats them as hypotheses—drafts of reality that need stress-testing by the people they’re meant to represent.
In practice, this means setting up iterative validation checkpoints where real participants, domain experts, or stakeholder panels review AI-generated focus group transcripts and flag what feels implausible, missing, or off-key. These reviews aren’t post-hoc cleanups; they’re built into the workflow. Each pass recalibrates the model, closing the gap between simulated consensus and lived complexity. When done well, this kind of auditing reframes AI from a storyteller into a listener—one that learns from correction, not repetition.
Human-in-the-loop research already shows what this can look like at scale. In AdaTest++, auditors working alongside AI discovered novel model failures in minutes, precisely because they weren’t deferential—they were curious.24 Their interventions reshaped the algorithm’s logic rather than simply diagnosing it. Similar approaches appear in climate policy modeling, where human experts adjust carbon-pricing simulations to ensure that projections align with political realities, social values, and local priorities.25 These iterative corrections may not guarantee precision, but they tether technical insight to moral ground and practical relevance.
Now imagine applying that same reflexive structure to synthetic qualitative research. Instead of a single model output, you get a dialogue between the machine’s first pass and the human community’s lived sense of truth. When auditors are embedded in the research design process to critique an AI’s rendering of their perspectives, it changes the tone of engagement entirely. Other sectors illustrate how this symbiosis might play out. In energy grid optimization, domain specialists collaborate with AI systems to fine-tune renewable integration parameters and respond to real-time disruptions.25 Their feedback doesn’t undermine the algorithm; it strengthens its adaptability and contextual intelligence. Translating that to synthetic focus groups, auditors would serve a similar function—spotting contextual mismatches, identifying cultural nuance, and retraining the model to recognize them before misinterpretation hardens into policy.
Behaviorally, this process leverages a simple mechanism: interruption. Automated systems—and the humans who use them—tend toward fluency bias, the sense that smooth outputs signal truth. Auditing introduces deliberate friction, reminding teams that confidence and correctness aren’t synonyms. It may slow analysis slightly, but it also safeguards credibility by ensuring that each conclusion withstands scrutiny rather than gliding on coherence.
For decision-makers, the payoff is twofold. First, behavioral auditing preserves inclusivity by making dissent part of the model’s learning process, not a threat to it. Second, it builds traceability—each revision documents where human judgment intervened, why, and with what effect. That record doesn’t just strengthen accountability; it helps future teams understand how consensus was built and when it began to drift.
Synthetic focus groups can’t think for us, but they can help us think better—if they’re kept under watch. Behavioral auditing turns AI-driven inference into a living process, where inclusion isn’t declared, it’s verified. After all, the best conversations—synthetic or not—are the ones that keep space for dissent.
Caveats to Consider
Embedding behavioral auditing into synthetic focus groups may sound elegant, but implementation exposes deeper limits. Participation isn’t just a methodological choice—it’s a function of trust, history, and power. Communities that’ve been repeatedly studied without seeing benefits may already mistrust institutional research. That skepticism can easily extend to AI-mediated dialogue. Recent work in digital health shows that patients who mistrust their human providers often rate healthcare AI systems as equally untrustworthy and unfair.26 The same dynamic could shape recruitment for synthetic focus groups in any domain where institutional credibility is low. Without visible reciprocity or shared ownership of outcomes, participation can feel like surveillance rather than collaboration.
Another challenge isn’t that data from low-resource settings don’t exist—it’s that they’re often too limited, inconsistent, or context-bound to sustain meaningful generalization. Local health surveys, community transcripts, or ethnographic data might capture rich nuance within a region but lack the scale needed for robust model training. When those smaller datasets get mixed with broader, demographically skewed samples, local context can disappear into statistical noise—and that’s where the cracks start to show. In that light, rolling out synthetic focus groups for some populations might be premature, however well-intentioned. AI can simulate diversity, but without enough, or the right kind of, data, it risks staging representation rather than earning it.
Even the auditing layer can potentially reproduce the hierarchies it’s meant to correct. External reviewers often come from the same professional strata—academia, policy, or tech—that designed the systems in the first place. If community-level or lay participation isn’t built into the audit loop, fairness remains procedural, not lived.
Still, these limitations can be managed. With clearer data provenance, trust-building engagement strategies, and review panels that privilege contextual knowledge, many of these gaps can be closed. Synthetic focus groups won’t ever replace lived dialogue, but they can inch closer to it if they’re built with equity as a design principle rather than an afterthought.
Where Human Perspective and AI Insights Shape Research, Together
Synthetic focus groups won’t replace traditional research, but they might change how it works. Across these two challenges—representation gaps and data quality—the pattern’s familiar: innovation races ahead, but the evidence may struggle to keep up. When the training data leaves people out, no algorithm can add them back in. And when behavioral audits are skipped, even the smartest systems can relearn the same blind spots we’ve been trying to correct for decades.
Still, there’s reason for optimism. With behavioral science built in from the start, synthetic research can become a complement, not a shortcut—one that helps teams test their assumptions, find weak spots early, and design with greater empathy. Auditing data for bias might feel tedious, but it’s what keeps the future of AI-powered research trustworthy.
Whether you’re planning a qualitative study to see how your next product lands or trying to close gaps in healthcare access, synthetic focus groups may help you reach further without losing sight of who’s being reached. At The Decision Lab, we help bridge those worlds. Our work blends behavioral science with design strategy so innovation doesn’t just move quickly—it moves wisely. As synthetic methods mature, focus groups included, they’ll need guardrails grounded in psychology and human behavior. Partner with us to explore what responsible, human-centered AI research might look like in your field—and help shape a future where technology learns to listen before it speaks.
Related TDL Articles
Focus Group
Curious about how focus groups came to shape modern research—and why they still matter today? In this article, we’ll explore their history, examine common challenges in how they’re used, and highlight why, despite their flaws, they remain one of the most effective tools for uncovering how people think, feel, and decide about an array of products and services.
Behavioral Science is WEIRD and This Should Concern Us…
We’ve highlighted how one major concern in using AI-driven datasets is that samples from marginalized populations often aren’t studied as extensively as others. Much of behavioral and psychological research still draws from WEIRD populations—Western, Educated, Industrialized, Rich, and Democratic. This article considers how the WEIRD problem plays out in real-world research and what it might take to design data systems that reflect the full complexity of human experience.
Sources
- Leung, F. H., & Savithiri, R. (2009). Spotlight on focus groups. Canadian Family Physician, 55(2), 218–219. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2642503/
- Nagle, B., & Williams, N. (2013). Methodology brief: Introduction to focus groups. Center for Assessment, Planning and Accountability.
- Collier, D., & Mahoney, J. (1996). Insights and pitfalls: Selection bias in qualitative research. World Politics, 49(1), 56–91. https://doi.org/10.1353/wp.1996.0023
- Solberg, S. (2025). A guide to recruiting busy participants in qualitative research: Twelve tips that will enhance your chances of success. International Journal of Qualitative Methods, 24, 16094069251381017. https://doi.org/10.1177/16094069251381017
- Newington, L., & Metcalfe, A. (2014). Factors influencing recruitment to research: Qualitative study of the experiences and perceptions of research teams. BMC Medical Research Methodology, 14(1), 10. https://doi.org/10.1186/1471-2288-14-10
- Keen, S., Lomeli-Rodriguez, M., & Joffe, H. (2022). From challenge to opportunity: Virtual qualitative research during COVID-19 and beyond. International Journal of Qualitative Methods, 21, 16094069221105075. https://doi.org/10.1177/16094069221105075
- Zhang, T., Zhang, X., Cools, R., & Simeone, A. (2024, September). Focus agent: LLM-powered virtual focus group. In Proceedings of the 24th ACM International Conference on Intelligent Virtual Agents (pp. 1–10).
- Brüggen, E., & Willems, P. (2009). A critical comparison of offline focus groups, online focus groups and e-Delphi. International Journal of Market Research, 51(3), 1–15. https://doi.org/10.2501/S1470785309200608
- Abrams, K. M., Wang, Z., Song, Y. J., & Galindo-Gonzalez, S. (2015). Data richness trade-offs between face-to-face, online audiovisual, and online text-only focus groups. Social Science Computer Review, 33(1), 80–96. https://doi.org/10.1177/0894439313519733
- Danry, V. M. (2023). AI enhanced reasoning: Augmenting human critical thinking with AI systems. Massachusetts Institute of Technology.
- Hinton, M., & Wagemans, J. H. (2023). How persuasive is AI-generated argumentation? An analysis of the quality of an argumentative text produced by the GPT-3 AI text generator. Argument & Computation, 14(1), 59–74. https://doi.org/10.3233/AAC-210026
- Erves, J. C., Mayo-Gamble, T. L., Malin-Fair, A., Boyer, A., Joosten, Y., Vaughn, Y. C., Sherden, L., Luther, P., Miller, S., & Wilkins, C. H. (2017). Needs, priorities, and recommendations for engaging underrepresented populations in clinical research: A community perspective. Journal of Community Health, 42(3), 472–480. https://doi.org/10.1007/s10900-016-0279-2
- Baird, K. L. (2020). The new NIH and FDA medical research policies: Targeting gender, promoting justice. Women, Medicine, Ethics and the Law, 237–271. https://doi.org/10.1215/03616878-24-3-531
- Pardhan, S., Sehmbi, T., Wijewickrama, R., Onumajuru, H., & Piyasena, M. P. (2025). Barriers and facilitators for engaging underrepresented ethnic minority populations in healthcare research: An umbrella review. International Journal for Equity in Health, 24(1), 1–16. https://doi.org/10.1186/s12939-025-02431-4
- Cao, S., Liu, Y., Song, W., Cui, Z., Lv, X., & Wan, J. (2019, November). Toward human-in-the-loop prohibited item detection in X-ray baggage images. In 2019 Chinese Automation Congress (CAC) (pp. 4360–4364). IEEE.
- Humans in the Loop. (n.d.). Your partner for high-quality AI data [Website]. https://humansintheloop.org/
- Shahbazi, N., Lin, Y., Asudeh, A., & Jagadish, H. V. (2022). A survey on techniques for identifying and resolving representation bias in data. CoRR, abs/2203.11852. https://arxiv.org/abs/2203.11852
- Gichoya, J. W., Thomas, K., Celi, L. A., Safdar, N., Banerjee, I., Banja, J. D., ... & Purkayastha, S. (2023). AI pitfalls and what not to do: Mitigating bias in AI. The British Journal of Radiology, 96(1150), 20230023. https://doi.org/10.1259/bjr.20230023
- Joseph, J. (2025). Algorithmic bias in public health AI: A silent threat to equity in low-resource settings. Frontiers in Public Health, 13, 1643180. https://doi.org/10.3389/fpubh.2025.1643180
- Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342
- Cronjé, H. T., Katsiferis, A., Elsenburg, L. K., Andersen, T. O., Rod, N. H., Nguyen, T. L., & Varga, T. V. (2023). Assessing racial bias in type 2 diabetes risk prediction algorithms. PLOS Global Public Health, 3, e0001556. https://doi.org/10.1371/journal.pgph.0001556
- Challen, R., Denny, J., Pitt, M., Gompels, L., Edwards, T., & Tsaneva-Atanasova, K. (2019). Artificial intelligence, bias and clinical safety. BMJ Quality & Safety, 28(3), 231–237. https://doi.org/10.1136/bmjqs-2018-008370
- Carter, L., & Liu, D. (2025). How was my performance? Exploring the role of anchoring bias in AI-assisted decision making. International Journal of Information Management, 82, 102875. https://doi.org/10.1016/j.ijinfomgt.2025.102875
- Rastogi, C., Tulio Ribeiro, M., King, N., Nori, H., & Amershi, S. (2023, August). Supporting human-AI collaboration in auditing LLMs with LLMs. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (pp. 913–926). https://doi.org/10.1145/3600211.3604712
- Debnath, R., Tkachenko, N., & Bhattacharyya, M. (2025). Enabling people-centric climate action using human-in-the-loop artificial intelligence: A review. Current Opinion in Behavioral Sciences, 61, 101482. https://doi.org/10.1016/j.cobeha.2025.101482
- Lee, M. K., & Rich, K. (2021, May). Who is included in human perceptions of AI?: Trust and perceived fairness around healthcare AI and cultural mistrust. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (pp. 1–14). https://doi.org/10.1145/3411764.3445570















