Synthetic Focus Groups: When AI-Simulated Consumers Work and When They Mislead

Are synthetic focus groups a breakthrough or a bias trap? Learn about the promise and pitfalls of using AI simulations to replace traditional consumer research.

The modern focus group was born out of wartime necessity. In the 1940s, sociologist Robert Merton began convening groups of listeners to evaluate government radio broadcasts, searching for the messages most likely to sustain morale and discourage defection. What he found was that the group setting itself was generative: participants built on each other's reactions, surfacing responses that no individual interview could have uncovered.1 

In the eighty-odd years since, focus groups have become a cornerstone of consumer research—and a remarkably expensive one at that. Recruiting panelists, renting facilities, and compensating participants for their time can cost tens of thousands of dollars for a single study. For smaller companies, or those working under tight release schedules, that investment can be prohibitive. 

Synthetic focus groups (SFGs) promise a way out. By training AI agents on data collected from real people, researchers can run consumer simulations at a fraction of the cost—and in a fraction of the time. The pitch is seductive: diverse, scalable, on-demand market testing for companies that have never been able to afford the real thing. Like most seductive pitches, however, it deserves some scrutiny.

Too good to be true?

The key question is whether or not SFGs give answers that reflect the attitudes of real people, and not just averages of populations. If a simulated agent can give a reasonably faithful representation of an individual’s taste profile, it can help us to understand how a product will play against the idiosyncratic preferences of real-life consumers.

A study conducted by Stanford and Google DeepMind compared the answers given by synthetic subjects with those given by their real counterparts.2 Researchers trained AI agents on individual humans. They then put both the humans and the agents through a battery of questions, including the Big Five Personality Inventory, several behavioral experiments, and a general social survey. They found that the AI agents gave the same answers as their human referents about 85% of the time.

On its own, this seems fairly useful. Real consumers are messy and irrational; an 85% accurate simulation that can be run on an unlimited basis is probably as directionally accurate as a real time-limited focus group. This route is less appealing, however, when we take into account the extent of the training necessary to produce these results. The researchers conducted two-hour interviews with each of their human subjects and used those to train the models. In effect, they had to impanel a real focus group in order to make a reliable synthetic one.

The most likely way for a technology like this to be implemented would be for a company to conduct thousands of interviews with real people, and then maintain a bank of synthetic subjects that market researchers can rent for SFGs. This eliminates the need to train new agents for every group. At the same time, it means that market researchers across the world will be testing their products on the same limited number of people.

Additionally, the errors made by models are not random. Synthetic users are likely to systematically diverge from real ones on certain subjects. A podcast sponsored by The Times newspaper asked a synthetic focus group which new features they’d like to see. The AI agents suggested that The Times add AI features such as explainers and summaries; however, the podcast found that its real listeners were deeply skeptical about AI add-ons.3

It’s also important for market researchers to note that the Stanford study was primarily designed to test the AI agents’ ability to answer questions about themselves. This does not mean that they grasp the humans’ preferences in real life. A real focus group panelist may love a proposed product name because it reminds them of a line in a film they enjoy. This may seem like noise, but it can be important information in the aggregate. Perhaps that line is widely enjoyed by people in the target demographic, providing a built-in positive predisposition towards the product. AI agents do not have this kind of history, and thus cannot capture this type of preference.

Ethical concerns

Synthetic focus groups are not without controversy. It’s worth taking a moment to consider what’s really happening when we use simulated users to diversify our studies. The AI agent may be pulling from an interview with a real person, but that interview isn’t comprehensive. Where there are gaps, the model will use its general knowledge to expand the areas it can speak to. It will pull from what it knows about people from the background it is told to represent. In other words, it will make generalizations. The line here between diversifying a dataset and perpetuating stereotypes is very thin. Synthetic users cannot have the particular experiences that make diverse perspectives so enriching. Instead, they substitute broad trends. 

For example, a study found that Indian users were more likely to use bland clichés to describe their preferences when writing essays with AI assistance than without. They named western foods like pizza and lasagna as their favorites at a higher rate, and when they talked about Indian food, they used vaguer language. Essays that were entirely human-written highlighted specific details: the smell of cardamom, the tang of curry leaves. AI-assisted essays would instead use phrases like “rich flavors and spices,” which might be found in Western cooking guides.4

There's a deeper philosophical problem here. Diverse representation in research isn't just about capturing statistical variance—it's about including voices that have historically been excluded from decisions that affect them. A synthetic agent standing in for a Kenyan grandmother or a first-generation college student doesn't amplify that person's perspective. It replaces it with an aggregate, laundered through a model trained predominantly on the kinds of text that reach the internet. The result can be the opposite of inclusion: a simulation of diversity that, in practice, reinforces the assumptions researchers brought to the exercise.

Privacy concerns compound these worries. In the Stanford model, AI agents are trained on two-hour interviews which are dense with personal information about tastes, anxieties, political leanings, and daily routines. The people who sit for these interviews may not fully understand how their data will be used, who else will have access to their synthetic double, or how long that double will remain in circulation. If a company builds a bank of synthetic subjects for other researchers to rent, the original participants are effectively licensing their psychological profiles to strangers, indefinitely, for purposes they never anticipated. Existing data protection frameworks were not designed with this use case in mind, and so far, the industry has moved faster than the regulations meant to govern it.

Getting it right

None of this means that synthetic focus groups lack value. The same qualities that limit their usefulness as a replacement for real research make them genuinely powerful as a supplement to it. When a product team wants to quickly stress-test a concept before committing to full development, an SFG can surface obvious problems cheaply and early. When a question requires a depth of engagement that would exhaust any real panelist—say, evaluating dozens of packaging variations or working through a complex pricing matrix—synthetic subjects can absorb the volume. Used this way, SFGs function less like focus groups and more like a sophisticated drafting tool: a way to sharpen the questions before you bring them to real people.

That sequencing matters. The Times example illustrates that the errors synthetic subjects make are not random, instead clustering around the kinds of lived, idiosyncratic experience that focus groups exist to surface. An SFG cannot tell you that a product name reminds your target demographic of a beloved film, or that a color scheme feels funereal to people who grew up in a particular region. It cannot replicate the way one panelist's offhand comment unlocks a feeling that five others didn't know they had. These are not incidental features of human research. They are the whole point.

The promise of synthetic focus groups is real but bounded. They can make market A)research faster, cheaper, and more accessible. What they cannot do is replace the unruly, expensive, irreducibly human process of actually asking people what they think. Merton's original insight—that the group setting itself generates knowledge—still holds. Synthetic subjects can help us prepare for that conversation. They cannot have it for us.

References

  1. Bloor, M., Frankland, J., Thomas, M., & Robson, K. (2001). Trends and uses of focus groups. In Trends and uses of focus groups (pp. 1-18). SAGE Publications Ltd, https://doi.org/10.4135/9781849209175.n1
  2. Park, J. S., Zou, C. Q., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., Willer, R., Liang, P., & Bernstein, M. S. (2024). Generative agent simulations of 1,000 people (arXiv:2411.10109). arXiv. https://doi.org/10.48550/arXiv.2411.10109
  3. Peterson, T. (2025, November 4). How The Times is using AI to model synthetic focus groups from human audiences. Digiday. https://digiday.com/media/how-the-times-is-using-ai-to-model-synthetic-focus-groups-from-human-audiences/
  4. Chayka, K. (2025, June 25). A. I. Is homogenizing our thoughts. The New Yorker. https://www.newyorker.com/culture/infinite-scroll/ai-is-homogenizing-our-thoughts

About us

We are the leading applied research & innovation consultancy

Our insights are leveraged by the most ambitious organizations

Image

I was blown away with their application and translation of behavioral science into practice. They took a very complex ecosystem and created a series of interventions using an innovative mix of the latest research and creative client co-creation. I was so impressed at the final product they created, which was hugely comprehensive despite the large scope of the client being of the world's most far-reaching and best known consumer brands. I'm excited to see what we can create together in the future.

Heather McKee

BEHAVIORAL SCIENTIST

GLOBAL COFFEEHOUSE CHAIN PROJECT

OUR CLIENT SUCCESS

$0M

Annual Revenue Increase

By launching a behavioral science practice at the core of the organization, we helped one of the largest insurers in North America realize $30M increase in annual revenue.

0%

Increase in Monthly Users

By redesigning North America's first national digital platform for mental health, we achieved a 52% lift in monthly users and an 83% improvement on clinical assessment.

0%

Reduction In Design Time

By designing a new process and getting buy-in from the C-Suite team, we helped one of the largest smartphone manufacturers in the world reduce software design time by 75%.

0%

Reduction in Client Drop-Off

By implementing targeted nudges based on proactive interventions, we reduced drop-off rates for 450,000 clients belonging to USA's oldest debt consolidation organizations by 46%

Read Next

Notes illustration

Eager to learn about how behavioral science can help your organization?