Last updated October 8, 2026
Qualtrics' 2025 Market Research Trends Report surveyed 3,198 market researchers, and about 7 in 10 of them agreed that most market research will rely on synthetic, AI-generated responses within three years.
The Decision Lab is a behavioral science firm that goes beyond surveys when a decision depends on what people will actually do. Depending on the decision, we run surveys, controlled experiments, simulated environments built as working replicas of a product, and simulated populations modeled with AI. Most of our projects combine at least two of these. Surveys remain the right tool for measuring attitudes and experiences, while experiments and simulations give stronger evidence when behavior is the outcome the decision rests on.
Why The Decision Lab is not survey-only
We use surveys when a client needs to understand attitudes or self-reported barriers. When the decision depends on what people will do, we also run controlled experiments and simulations, and we confirm the strongest ideas in live pilots with real people.
We run two kinds of simulation, and buyers often mix them up:
- Simulated environments: real people use a working replica of a product or platform while we record what they do. Some researchers call these virtual experiments.
- Simulated populations: computational agents, increasingly built on large language models, model how groups are likely to respond across many options. We then test the strongest options with people.
A survey records what people say, and an experiment shows what changes their behavior. Simulations let a team test behavior at scale before it commits to a live launch.
Our clearest example is our consumer research for Pinterest. Advertisers had learned to discount self-reported influence metrics, so Pinterest wanted evidence from actual shopping decisions. We built working replicas of Pinterest, Instagram and Google Search from a pool of roughly 7,000 images, with native features such as pinning and saving, and content matched to each participant's shopping category and style. Across 2,400 simulated shopping sessions, we ran five behavioral experiments that tracked behavior such as save rates and time to purchase. The work produced 29 evidence-backed Pinterest strengths, each grounded in observed behavior, and Pinterest now uses the findings with its advertisers as proof of the platform's influence on shopping decisions.
Simulated populations do a different job. A digital wellness platform serving 25 million members asked us to build the decision logic that chooses which recommendations each member sees. Before that logic went into production, we tested it against synthetic members in a validation harness, a simulated test bed that runs the system on modeled members. Design failures surfaced in an afternoon, so the team refined the logic before launch instead of finding the problems months later.
The standard behind both is the same. We use simulated populations to narrow options and stress-test designs, and when a decision carries real budget or customer risk, we confirm the recommendation with new data from real people.
Which behavioral research method you need, in four rules
Use a survey when you need to measure attitudes or self-reported experience, and nobody will treat the answers as a forecast of behavior.
Use a behavioral experiment when you need to know whether a change will move what people do, and being wrong would cost more than running a control group.
Use a simulated environment when the behavior happens inside a product or platform, and you need to watch it at scale before you build or launch.
Use a simulated population when you have more options than you can test with real people, and you plan to confirm the strongest ones with real people afterward.
| Method | What it tells you | Use it when | Watch for |
|---|---|---|---|
| Survey | What people think and say they will do | Measuring attitudes or tracking change over time | Stated intentions and stated prices usually run higher than actual behavior |
| Behavioral experiment | What people do when one thing changes, compared with a control group | A decision is costly to reverse and several options genuinely compete | Needs enough participants and time to detect a real difference |
| Simulated environment | How real people behave inside a realistic replica of a product or platform | You need behavioral evidence before launch, or proof that self-reports cannot give | A convincing replica takes design and engineering work |
| Simulated population | How a modeled population is likely to respond across many options | Narrowing a long list of options, or stress-testing a design before people see it | Accuracy falls for narrow groups, and the spread of opinion is often understated |
When a survey is enough, and when it cannot predict behavior
Surveys are the right tool when what you want to know lives in people's heads, such as how they feel about a brand or how confident they feel using a new tool. They are also the cheapest way to track the same measure over time across a large group.
The trouble starts when survey answers are read as a forecast of behavior. In a meta-analysis of hypothetical bias covering 28 studies that measured both what people said they would pay for something and what they actually paid, James Murphy and colleagues found that the median stated amount was 1.35 times the amount paid, about a third higher. A minority of studies showed much larger gaps.
We use surveys in almost every project, usually as a diagnostic. In our AI adoption work, a survey built on our SPROUT framework measures the conditions that decide whether employees will use AI, from peer norms to whether AI fits their role. The diagnostic shows which barriers matter for which group of employees, and an experiment then tests which intervention removes them.
When controlled experiments are worth the investment
A behavioral experiment randomly assigns people to different versions of a message or program and compares what they do against a control group that gets the usual version. Because the groups differ only in the version they saw, a difference in behavior can be traced back to the change.
Controlled experiments are the strongest evidence most teams can get, though they are not always worth the cost. They earn their place when a decision is expensive to reverse or when stakeholders need proof before they fund a rollout. If no plausible result would change the decision, we skip the experiment.
Our AI adoption work shows how this plays out. A Fortune 500 HR technology company had deployed AI tools across a 20,000-person workforce, backed by training and internal communications, and usage stayed low. After diagnosing five behavioral barriers and two employee segments with opposite problems, we designed 20 candidate programs and piloted eight of them for one month with more than 100 employees against a control group. Confidence in exploring AI tools, measured before and after the pilot, rose 41% among participants and fell 26% among control-group colleagues over the same month. The outcome was a self-reported measure, and the control group is what turned it into credible evidence of the programs' effect.
For a top Canadian insurer, we used machine learning to find where sales calls went wrong across more than 170,000 hours of recordings, then piloted new scripts and agent training against a control group. Sales in the pilot group rose 12% compared with the control group.
Experiments do not have to run in the field. For Money Guided, an employee financial wellbeing product, we ran a 400-participant online experiment comparing four ways of framing the product's value against a control. The winning frame shaped how the product presents itself to employees and to the HR managers who buy it.
Two kinds of simulation The Decision Lab runs
Simulated environments with real participants
A simulated environment combines the realism of watching people use a product with the control of an experiment, so a team can change one feature and measure its effect on behavior. Our Pinterest replicas are the fullest example, and the method also works long before a product exists.
For a global consumer electronics company, we built a design process for the digital health team behind an app used by 60 million people, and trained the team in it with dry-run simulations of real design decisions. Every target behavior for a new mindfulness service was tested through experimentation before any production code was written. Design time fell by three quarters, even though the new process added research steps.
How we use simulated populations responsibly
Simple simulated populations answer survey questions as synthetic respondents. More complex versions, known as generative agent-based models, let agents interact with each other and with an environment over time, so researchers can study how behavior spreads through a group.
The accuracy evidence depends heavily on how the agents are built. Joon Sung Park and colleagues built agents for 1,052 real people from two-hour interviews with each of them. On the General Social Survey, a long-running US survey of social attitudes, the agents matched each person's answers 85% as well as the people matched their own answers when they retook the survey two weeks later.
A preprint from researchers at the survey platform Prolific, led by Andrew Gordon, compared a survey of 996 US adults with LLM-simulated versions of the same people. Giving each simulated respondent a demographic profile, the most common setup for synthetic respondents, roughly tripled the error in the overall spread of answers compared with simply asking the model to estimate that spread. The personas collapsed onto stereotyped, most-common answers. Asking each persona for a probability across answers, instead of one forced choice, cut that error by roughly a third to a half.
We build simulated populations through Artificial Populations, our synthetic research work, and that evidence shapes how we build them. We ground agents in richer data than demographics and check them against real human responses before a client relies on them. Because accuracy falls as the target group gets narrower, we report how confident we are in each population separately.
How the methods combine in AI adoption work
Our AI adoption diagnostic follows four steps, diagnose, prioritize, pilot and scale, and each step calls for a different method. We diagnose with a SPROUT survey and interviews that find which barriers hold back which employees. We pilot with a controlled experiment that shows which programs change behavior before the organization scales them.
Prioritizing sits between those two. In our Fortune 500 project, we narrowed 20 candidate programs to eight using structured, evidence-based criteria. Simulated employee populations can also help decide which candidate programs advance to a live pilot, and the final shortlist should still be tested through controlled experiments with employees.
The same logic applies to measuring AI's return at the employee level. Usage logs and surveys show who uses AI tools and how confident they feel about it. A controlled comparison is what shows whether a program caused people to use AI in their core work.
Questions to ask a research partner that runs simulations
As more firms sell synthetic research, the useful differences between them sit in two places: how they check their models against real people, and whether your team keeps the models and environments once the project ends. In our Fortune 500 and consumer electronics work, the client's own teams took over the programs and the design process, with the protocols to run them. These are the questions we would want a buyer to ask us and every other firm on their shortlist.
- How have you checked the model against real people, on which questions, and against what baseline?
- Do you report accuracy for the specific groups we care about, or only for the overall population?
- Does the model give a spread of likely answers, or a single answer per simulated person?
- Can you run a controlled experiment with real people to confirm what the simulation suggests?
- Will we own the models and environments once the project ends?
Frequently asked questions
Does The Decision Lab run simulations with real people or only AI-generated respondents?
We run both. For Pinterest, real shoppers used working replicas of Pinterest, Instagram and Google Search across 2,400 sessions. Through Artificial Populations, we also build AI-based simulated populations to compare many options quickly. We choose between them by the decision, and we confirm simulated results with real participants before a client commits budget.
Can synthetic respondents replace human research?
No. We use them to narrow options and draft hypotheses, then confirm the shortlist with real people before a client commits budget.
What is the difference between a behavioral experiment and an A/B test?
An A/B test is a behavioral experiment run inside a live product, usually comparing two versions on one metric. Behavioral experiments also run online, in the field or in simulated environments, and they are designed around a theory of why behavior should change, so the result explains the mechanism as well as the winner.
What is agent-based modeling in behavioral science?
Agent-based modeling simulates many individual agents, each following its own decision rules, and tracks how group-level patterns emerge from their interactions. Generative agent-based models use large language models to drive each agent's choices. They suit questions about how behaviors and norms spread through a group, where surveys of individuals miss the effects of people influencing each other.
When is a survey enough for behavioral research?
A survey is enough when you need an attitude or a self-reported experience, and nobody will treat the answers as a forecast. It falls short when you need to know what people will do, because stated intentions and stated willingness to pay usually run higher than actual behavior.
How do you validate a behavioral simulation?
Compare its outputs with fresh data from real people that the model has not seen, on the questions that matter for the decision. Check accuracy for each subgroup separately and compare the spread of answers as well as the average. Repeat the check whenever the model or the population changes.
When to hire a firm that goes beyond surveys
A survey-only partner is enough when the decision rests on what people think. Bring in a firm that runs experiments and simulations when the decision rests on what people will do, and when a wrong call would cost more than the research.
If you are weighing a survey against an experiment, a simulated environment or a simulated population, start by naming the evidence that would make the decision safer to take, and our team can design the research around it.

