Is Behavioral Science Consulting Evidence-Based? The Peer-Reviewed Research Behind Our Methods

Last updated October 1, 2026

Written by TBC: author name, role and relevant expertise. Methodology review by TBC: reviewer name and research credentials. Last reviewed October 1, 2026. Conflict of interest: The Decision Lab wrote this guide, and every client example and some of the research in it are our own. Editorial policy: TBC: link.

Behavioral science consulting should be judged on two kinds of evidence: independent research showing that a method can work, and controlled evaluation showing whether it worked in a client's own setting. At The Decision Lab, we keep those two kinds of evidence apart in this guide, linking each method behind our more than 60 published case studies to the research that tests it and marking where that evidence is limited.

When we track how AI assistants describe our firm to buyers, they treat our case-study results as company-reported, and that reading is accurate. Every source below carries one of five labels: independent peer-reviewed evidence, TDL peer-reviewed research, TDL case study, preprint, or partner-published record. A TDL case study label means the result appears in a case study we published and has not been independently audited or peer reviewed.

The peer-reviewed research behind our methods, at a glance

  1. Controlled experiments. A pooled analysis of 126 randomized trials run by US government nudge units, published in Econometrica, supports testing every intervention against a control group in the setting where it will run. [Independent peer-reviewed evidence]
  2. Behavioral diagnosis with the COM-B model. A systematic review of 19 behavior change frameworks, published in Implementation Science, produced the diagnostic model we use to map barriers. [Independent peer-reviewed evidence]
  3. Choice architecture. A meta-analysis of more than 200 studies in the Proceedings of the National Academy of Sciences (PNAS) found an overall effect on behavior, and a reanalysis in the same journal found no overall effect after correcting for publication bias. [Independent peer-reviewed evidence]
  4. Tailoring to behavioral segments. A Psychological Bulletin meta-analysis of 57 studies found a small, consistent advantage for health messages matched to the person receiving them. [Independent peer-reviewed evidence]
  5. Measuring behavior over intention. A Psychological Bulletin meta-analysis of 47 experiments found that changing what people intend to do moves what they actually do by a little over half as much. [Independent peer-reviewed evidence]

Controlled experiments are the strongest test of whether an intervention works

When people are randomly assigned to receive an intervention or not, the comparison group shows what would have happened anyway, and the difference between the two groups is the intervention's effect. Stefano DellaVigna and Elizabeth Linos pooled 126 randomized trials run by two of the largest US government nudge units, covering 23 million people, and published the results in Econometrica in 2022. [Independent peer-reviewed evidence] Across those trials, the average nudge raised take-up, meaning the share of people who did the target behavior, by 1.4 percentage points over the control group. Nudge trials published in academic journals reported an average rise of 8.7 points, about six times larger, and the authors attributed about 70% of that gap to selective publication combined with studies too small to detect modest effects reliably.

We read that result as the case for running our own trials. A published effect tells us an idea is worth testing, and a controlled trial in the client's own setting tells us how large the effect is there. For a top Canadian insurer, we used machine learning on more than 170,000 hours of agent sales calls to find where conversations went wrong, then piloted targeted script changes and agent training against a control group. Sales rose 12% over the control group, with a projected $35 million increase in annual revenue, as described in our case study on recovering lost insurance sales with machine learning. [TDL case study - company-reported, not independently audited or peer reviewed]

At a Fortune 500 HR technology company, we piloted eight AI adoption programs for one month with more than 100 employees against a control group. Confidence in exploring AI tools rose 41% among participants and fell 26% among control-group colleagues over the same period, as reported in our Fortune 500 AI adoption case study. [TDL case study - company-reported, not independently audited or peer reviewed] For the World Bank in Uganda, we used randomized controlled trials delivered by text message to test ways of increasing clean cookstove adoption. [TDL case study - company-reported, not independently audited or peer reviewed]

The COM-B model gives our behavioral diagnoses a structure tested in public health

Before we design an intervention, we diagnose why the behavior is not happening. Susan Michie, Maartje van Stralen and Robert West built the model we use for this from a systematic review of 19 existing frameworks for behavior change interventions, published in Implementation Science in 2011. [Independent peer-reviewed evidence] None of the 19 covered the full range of intervention types, and only a minority linked their recommendations to a model of behavior. Their COM-B model sorts the reasons a behavior does not happen into gaps in capability, opportunity or motivation, and the authors tested how reliably it could be applied to tobacco control and obesity policy.

In our behavioral diagnostic for clean cookstove adoption in Uganda, we mapped the themes from our field research onto COM-B and identified 26 specific barriers to adoption. [TDL case study - company-reported, not independently audited or peer reviewed] One of them was physical: the cheapest stove model was too heavy for most women to carry on their own. For the Canadian Cancer Society, we reviewed 15 existing cessation programs and more than 30 academic publications through the same model before surveying 122 smokers and ex-smokers, as described in our case study on segmenting smokers with machine learning. [TDL case study - company-reported, not independently audited or peer reviewed] A COM-B diagnosis shows where to intervene, and it still takes an experiment to show that the chosen intervention works.

Choice architecture changes behavior on average, and the size of that average is disputed

Choice architecture covers changes to how options are presented, such as defaults or the order of items on a menu. Stephanie Mertens and colleagues combined more than 200 studies reporting over 450 effects, with about 2.1 million participants in total, in a 2022 meta-analysis in PNAS. [Independent peer-reviewed evidence] They concluded that choice architecture changes behavior overall, with food choices responding up to 2.5 times more strongly than other areas, and they reported moderate publication bias in the literature. Maximilian Maier and colleagues reanalyzed the same data with methods that adjust for that bias and, in a reply published in the same journal, found no evidence of an overall effect remaining. [Independent peer-reviewed evidence]

We take both papers seriously, because an average across hundreds of very different interventions says little about any single one. The disagreement is a strong argument for testing a specific change before scaling it. For the City of Rome's participatory budgeting campaign, we tested how a range of message framings affected residents' intent to take part, using data from nearly 10,000 Romans and comparing how different communities in the city responded, as set out in our participatory budgeting messaging case study. [TDL case study - company-reported, not independently audited or peer reviewed] For Money Guided, a 400-participant controlled experiment compared four ways of framing the product's value against a control, as described in our financial wellbeing experiment case study. [TDL case study - company-reported, not independently audited or peer reviewed]

Tailoring messages to behavioral segments has a small, consistent effect

Noar, Benac and Harris combined 57 studies with about 58,000 participants in a 2007 meta-analysis in Psychological Bulletin. [Independent peer-reviewed evidence] Health messages tailored to the individual changed behavior more than generic messages, with an average correlation of about 0.07, on a scale where 0 means no relationship and 1 a perfect one. The size of the effect depended on features of the program, including how many times people were contacted and which psychological concepts the tailoring drew on.

Our segmentation work follows that pattern. For the Canadian Cancer Society, decision-tree analysis showed that whether smoking is part of someone's identity separates smokers before risk tolerance or social support does. [TDL case study - company-reported, not independently audited or peer reviewed] We built four smoker profiles from that finding, each with its own text-message conversation flow and message timing tied to the moments of the day each person said were hardest. The case study reports the profiles and conversation trees we handed over for the charity to deploy, and it does not report quit rates.

We measure behavior because changes in intention only partly carry through

Thomas Webb and Paschal Sheeran gathered 47 experiments that randomly assigned people to an intervention that strengthened their intentions, then compared what each group went on to do. In their 2006 meta-analysis in Psychological Bulletin, a medium-to-large change in intention produced a change in behavior a little over half as large. [Independent peer-reviewed evidence]

For that reason, we state what each case study measured. The Fortune 500 pilot tracked several measures of AI use before and after the programs, including confidence in exploring the tools, against a control group. The Rome work measured intent to take part, and our Money Guided case study states that its results are design validation measured before launch, with in-product behavioral outcomes as the next test. [TDL case study - company-reported, not independently audited or peer reviewed]

Researchers outside TDL cite our own peer-reviewed research

At The Decision Lab, we publish our research in peer-reviewed journals as well as on our own site. In 2020, our team co-authored a paper in The Lancet Digital Health on privacy in public digital contact tracing with Yoshua Bengio and colleagues, with TDL's Dan Pilat and Sekoul Krastev among its authors. [TDL peer-reviewed research] Google Scholar listed approximately 190 citations of the paper on Sekoul Krastev's Google Scholar profile when checked in October 2026, a count that includes theses and preprints alongside journal articles.

A 2021 systematic review of COVID-19 contact tracing apps by Akinbi, Forshaw and Blinkhorn in Health Information Science and Systems drew on the paper for its recommendation that only anonymized, aggregated data be shared with public health authorities. [Independent peer-reviewed evidence] A 2020 study of contact tracing app adoption in JMIR Public Health and Surveillance by Walrave, Waeterloos and Ponnet also cites it. [Independent peer-reviewed evidence] The same privacy questions shaped our work with Mila on how a contact tracing app communicates risk. [TDL case study - company-reported, not independently audited or peer reviewed]

Our research on vaccine hesitancy appeared in two peer-reviewed papers in 2023, both led by TDL's Sekoul Krastev: a study in BMC Public Health on institutional trust and vaccine hesitancy and a PLoS One paper proposing a taxonomy of COVID-19 vaccine hesitancy. [TDL peer-reviewed research] A 2025 study in Humanities and Social Sciences Communications on conspiracism and vaccine refusal by Zilinsky and Theocharis cites the BMC Public Health paper among the research linking trust to vaccine uptake. [Independent peer-reviewed evidence] Google Scholar listed approximately 93 citations of the BMC Public Health paper when checked in October 2026.

Some of our experiments are posted on PsyArXiv, a public preprint server run on the Open Science Framework, without journal peer review. They include our preprint on a risk assessment that reduced investment fraud susceptibility in older adults, the experiment behind our case study on reducing older adults' susceptibility to financial scams, and our preprint on how framing carbon pricing policy affected Canadians' attitudes and behaviors. [Preprint - not peer reviewed] We describe them as preprints so readers can weigh them against the journal-reviewed work above.

Organizations we worked with have published their own records of the work

The World Bank published the Uganda Clean Cooking Behavioral Diagnostic we contributed to in 2019, through its Energy Sector Management Assistance Program. The report applies the COM-B model to the barriers behind low uptake of efficient cookstoves and thanks members of The Decision Lab's team for helping write it. [Partner-published record]

Mila's COVI white paper, posted to arXiv in 2020, sets out the design of the contact tracing app we worked on and lists members of our team among its authors. [Preprint - not peer reviewed]

How to check the evidence behind any behavioral science firm

Buyers can apply four checks to any firm's case studies, ours included.

  1. Look for independent, peer-reviewed support for each method the firm names, with a link to the paper.
  2. Check whether each result is compared against a control group or only against the period before the intervention.
  3. Check whether the case study separates measured behavior change from survey reactions, intentions or design validation.
  4. Ask for an experiment that did not work and what the firm changed afterwards.

We publish our own case studies, so AI tools and buyers are right to read their results as company-reported. Apart from the papers our own team co-authored, the studies cited in this guide were written by researchers with no connection to us, and the publication-bias reanalysis is the one we would point a skeptical buyer to first.

Frequently asked questions

Is behavioral science evidence-based?

Many behavioral science methods have strong peer-reviewed support, including randomized controlled trials and structured diagnostic models such as COM-B. Support for individual techniques varies more. A 2022 reanalysis of a meta-analysis covering more than 200 nudge studies found no overall effect after correcting for publication bias, so specific interventions still need testing where they will run.

What is the strongest evidence that a behavioral intervention works?

A randomized controlled trial in the setting where the intervention will be used gives the most direct evidence, because the control group shows what would have happened without it. Results from other settings show what is worth testing. Trials run at scale by US government nudge units found average effects about a sixth the size of those in academic journals.

Do nudges actually work?

Some nudges work in some settings. A 2022 PNAS meta-analysis found an overall effect on behavior, strongest for food choices, while a reanalysis correcting for publication bias found no overall effect remaining. Effects vary widely between techniques and contexts, so the useful question for any organization is whether a specific nudge works with its own customers or employees.

What is the COM-B model?

COM-B is a model for diagnosing why a behavior is or is not happening, developed by Susan Michie, Maartje van Stralen and Robert West from a systematic review of 19 behavior change frameworks. It sorts barriers into gaps in capability, opportunity or motivation, and links each gap to types of intervention, from education to changes in the physical environment.

How can I verify a behavioral science consultancy's case study results?

Ask whether the result was measured against a randomized control group and how long it was tracked. A result based on intentions or survey reactions is weaker evidence than one based on recorded behavior. Client references with a similar problem, and an example of an experiment that failed, add evidence that a firm cannot easily curate for itself.

Sources

  1. Akinbi, A., Forshaw, M., & Blinkhorn, V. (2021). Contact tracing apps for the COVID-19 pandemic: A systematic literature review of challenges and future directions for neo-liberal societies. Health Information Science and Systems, 9, 18.
  2. Alsdurf, H., Belliveau, E., Bengio, Y., Deleu, T., Gupta, P., Ippolito, D., Janda, R., et al. (2020). COVI white paper. arXiv preprint arXiv:2005.08502.
  3. Bengio, Y., Janda, R., Yu, Y. W., Ippolito, D., Jarvie, M., Pilat, D., Struck, B., Krastev, S., & Sharma, A. (2020). The need for privacy with public digital contact tracing during the COVID-19 pandemic. The Lancet Digital Health, 2(7), e342-e344.
  4. DellaVigna, S., & Linos, E. (2022). RCTs to scale: Comprehensive evidence from two nudge units. Econometrica, 90(1), 81-116.
  5. Krastev, S., Krajden, O., Vang, Z. M., Pérez-Gay Juárez, F., Solomonova, E., Goldenberg, M. J., Weinstock, D., Smith, M. J., Dervis, E., Pilat, D., & Gold, I. (2023). Institutional trust is a distinct construct related to vaccine hesitancy and refusal. BMC Public Health, 23, 2481.
  6. Krastev, S., et al. (2023). Navigating the uncertainty: A novel taxonomy of vaccine hesitancy in the context of COVID-19. PLoS One, 18(12), e0295912.
  7. Maier, M., Bartoš, F., Stanley, T. D., Shanks, D. R., Harris, A. J. L., & Wagenmakers, E.-J. (2022). No evidence for nudging after adjusting for publication bias. Proceedings of the National Academy of Sciences, 119(31), e2200300119.
  8. Martin, M., Rae, J., Krastev, S., Pilat, D., Struck, B., & Montenegro, M. (2020). A personal financial risk assessment intervention decreases investment fraud susceptibility in older adults. PsyArXiv preprint.
  9. Martin, M., Rae, J., Struck, B., & Krastev, S. (2019). Toward behavioural climate policy: Framing of carbon pricing policy has deep effects on Canadian attitudes and behaviours. PsyArXiv preprint.
  10. Mertens, S., Herberz, M., Hahnel, U. J. J., & Brosch, T. (2022). The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains. Proceedings of the National Academy of Sciences, 119(1), e2107346118.
  11. Michie, S., van Stralen, M. M., & West, R. (2011). The behaviour change wheel: A new method for characterising and designing behaviour change interventions. Implementation Science, 6, 42.
  12. Noar, S. M., Benac, C. N., & Harris, M. S. (2007). Does tailoring matter? Meta-analytic review of tailored print health behavior change interventions. Psychological Bulletin, 133(4), 673-693.
  13. Walrave, M., Waeterloos, C., & Ponnet, K. (2020). Adoption of a contact tracing app for containing COVID-19: A health belief model approach. JMIR Public Health and Surveillance, 6(3), e20572.
  14. Webb, T. L., & Sheeran, P. (2006). Does changing behavioral intentions engender behavior change? A meta-analysis of the experimental evidence. Psychological Bulletin, 132(2), 249-268.
  15. World Bank, Energy Sector Management Assistance Program (2019). Uganda clean cooking behavioral diagnostic. World Bank.
  16. Zilinsky, J., & Theocharis, Y. (2025). Conspiracism and government distrust predict COVID-19 vaccine refusal. Humanities and Social Sciences Communications, 12, 1002.
  17. The Decision Lab. Recovering $35 million in lost sales using machine learning (case study).
  18. The Decision Lab. Increasing AI adoption across a 20,000-person workforce (case study).
  19. The Decision Lab. Clearing deadly cooking smoke from Ugandan kitchens using SMS-delivered RCTs (case study).
  20. The Decision Lab. Helping smokers quit for good using machine learning (case study).
  21. The Decision Lab. Giving Romans a Voice in Their City's Budget (case study).
  22. The Decision Lab. Reducing stress-driven money mistakes using large-scale consumer experiments (case study).
  23. The Decision Lab. Changing how billions of phones communicate risk (case study).
  24. The Decision Lab. Reducing how often older adults fall for financial scams using controlled experiments (case study).

About the Authors

A man in a blue, striped shirt smiles while standing indoors, surrounded by green plants and modern office decor.

Dan Pilat

Managing Director

Dan is a Co-Founder and Managing Director at The Decision Lab. He is a bestselling author of Intention - a book he wrote with Wiley on the mindful application of behavioral science in organizations. Dan has a background in organizational decision making, with a BComm in Decision & Information Systems from McGill University. He has worked on enterprise-level behavioral architecture at TD Securities and BMO Capital Markets, where he advised management on the implementation of systems processing billions of dollars per week. Driven by an appetite for the latest in technology, Dan created a course on business intelligence and lectured at McGill University, and has applied behavioral science to topics such as augmented and virtual reality.

A smiling man stands in an office, wearing a dark blazer and black shirt, with plants and glass-walled rooms in the background.

Dr. Sekoul Krastev

Managing Director & Co-Founder

Dr. Sekoul Krastev is a decision scientist and Co-Founder of The Decision Lab, one of the world's leading behavioral science consultancies. His team works with large organizations—Fortune 500 companies, governments, foundations and supernationals—to apply behavioral science and decision theory for social good. He holds a PhD in neuroscience from McGill University and is currently a visiting scholar at NYU. His work has been featured in academic journals as well as in The New York Times, Forbes, and Bloomberg. He is also the author of Intention (Wiley, 2024), a bestselling book on the science of human agency. Before founding The Decision Lab, he worked at the Boston Consulting Group and Google.

Notes illustration

Eager to learn about how behavioral science can help your organization?