Last updated September 28, 2026
Across 74 nudge interventions published in academic journals, the average rise in take-up, meaning the share of people who did the target behavior such as enrolling in a program, was 8.7 percentage points over a control group. When Stefano DellaVigna and Elizabeth Linos pooled 126 randomized trials run at scale by two US government nudge units, covering 23 million people, the average rise over their control groups was 1.4 points, about a sixth as large.
Choosing a behavioral science consulting firm mostly comes down to whether the firm can close that gap in your setting, by testing ideas with your own customers or employees and then staying involved until the winning version is built and running. At The Decision Lab, we combine behavioral research and experimentation with product design and implementation, and the ten questions below are the ones we would want any buyer to ask us and every other firm on their shortlist.
A five-point checklist for choosing a behavioral science consulting firm
A strong firm should be able to show you five things:
- It runs its own experiments and measures results against a control group.
- It tests ideas with your own customers or employees, and treats published research as a starting point.
- It can design and build the intervention as well as recommend it.
- Your team will own the methods and assets once the engagement ends.
- It can show measured outcomes, including experiments that did not work.
The 10 questions at a glance
- Academic rigor: Who on your team is trained in research methods, and will they design our study?
- Experimentation: Which experiments have you designed and run yourselves, from start to finish?
- Primary evidence: What do you do when the published research does not cover our situation?
- Causal measurement: How will we know your intervention caused the result?
- Domain expertise: Which past project most resembles ours, and what are you assuming about our case?
- Implementation and design: Who designs and builds what you recommend?
- Scalability: What do you expect to happen to the effect when we roll this out to everyone?
- Internal teams: What will our people be able to run on their own when you leave?
- Measurable outcomes: What changed for past clients, and compared with what?
- Uncertainty: What are you unsure about, and what result would change your plan?
1. Academic rigor
Question to ask: Who on your team is trained in research methods, and will they design our study?
Academic rigor in consulting shows up in method choices, such as committing to an outcome before any data comes in and recruiting enough participants to detect an effect of the size that would matter to the business. Credentials on a team page tell you less than whether those people will design your study, so ask to meet the researchers and have one of them walk through a past study and explain the sample size they chose.
Rigor also means reading the literature critically. In 2022, Stephanie Mertens and colleagues published a meta-analysis of nudging in the Proceedings of the National Academy of Sciences (PNAS) reporting that nudges change behavior with a small-to-medium average effect. Later that year, Maximilian Maier and colleagues published a reanalysis in the same journal that corrected for publication bias, the tendency for positive results to reach print more often than null ones, and found no remaining evidence for an average effect of nudging. The debate continues, partly because averaging across very different interventions hides how much they differ, and a rigorous firm can explain what it means for the specific change you are considering.
At The Decision Lab, we start most engagements with a structured review of the research and treat it as a set of hypotheses to test against data from the client's own customers or employees.
2. Experience running real experiments
Question to ask: Which experiments have you designed and run yourselves, from start to finish?
Many firms describe their work as experimental, but fewer can show a trial they designed, randomized, ran and analyzed themselves. Ask for two recent examples and, for each one, how participants were assigned to groups and what the firm did when a result came back smaller than expected.
Experiments come in different forms. Online experiments are fast and inexpensive, and they suit early questions such as which of several messages people understand best. Field trials take longer but measure what people actually do in the real setting, and product A/B tests sit somewhere between the two. A capable firm will explain which form fits your question and what each one cannot tell you.
Case example: financial wellbeing product
- Client type: A company building an employee financial wellbeing product (Money Guided).
- Problem: The product needed evidence on two audiences at once, the employees who use it and the HR managers who buy it.
- Method: Surveys of 533 UK employees and 264 HR managers, followed by a controlled experiment with 400 participants testing four ways of framing the product's value against a control group.
- Measured result: The experiment identified a winning frame. These are design-validation results measured before launch, and in-product behavior has not yet been measured.
- What changed in implementation: The features we specified went into Money Guided's first product release, and the winning frame shapes how the product presents its value.
3. Primary evidence versus literature alone
Question to ask: What do you do when the published research does not cover our situation?
A literature review tells you what worked for other people, in other places, often years ago. It is a necessary starting point, and the nudge-unit findings above show how much effects can shrink between published studies and real rollouts. Ask what share of a typical engagement's recommendations comes from the firm's own data, and what it does when the research is silent on your population.
Case example: sonic branding experiment
- Client type: A sonic branding agency (Made Music Studio).
- Problem: The agency wanted to know how the sound customers hear while waiting affects their experience, and our review found the link between music and perceived waiting time was not yet well understood.
- Method: A controlled experiment with nearly 200 participants comparing seven sound conditions, from ambient music to silence.
- Measured result: Participants felt time passed most quickly with ambient music.
- What changed in implementation: The project was designed to produce original evidence for the agency, and its deployment is outside the scope of the published case study.
Case example: clean cookstove adoption in Uganda
- Client type: An international development institution (the World Bank), working in Uganda.
- Problem: Clean cookstoves costing under $10 could cut indoor pollution by up to 90%, yet fewer than 1 in 20 Ugandan households used one as of 2018.
- Method: Field research, with the findings mapped onto COM-B, a widely used behavior change model.
- Measured result: None yet at this stage. The diagnostic identified 26 specific barriers, including one no general review would have predicted: the cheapest model was too heavy for most women to carry on their own.
- What changed in implementation: A later phase of the project tested interventions through randomized trials delivered by text message.
4. Causal measurement
Question to ask: How will we know your intervention caused the result?
Causal measurement compares people who received a change with a similar group who did not, so the difference can be credited to the change and not to seasonality or a campaign running at the same time. A randomized controlled trial, where chance decides who gets the change, gives the cleanest answer. When randomizing is not possible, staggered rollouts and carefully matched comparison groups can still work, with weaker guarantees that a good firm will spell out.
Ask each firm to name its comparison group and its primary outcome before the work starts, along with the smallest effect it would treat as worth acting on. A plan built on before-and-after numbers with no comparison group cannot separate the intervention from everything else that changed.
Case example: Canadian insurer
- Client type: One of North America's largest insurers, based in Canada.
- Problem: Sales were being lost inside agent calls, in ways no dashboard captured.
- Method: Analysis of more than 170,000 hours of sales calls, using process mapping and machine learning to find where calls went wrong, then a pilot of targeted script changes and agent training against a control group.
- Measured result: Sales rose 12% in the pilot group relative to the control group, with a projected $35 million increase in annual revenue.
- What changed in implementation: We also designed a permanent in-house behavioral science team, and the insurer now runs that team itself.
5. Domain expertise
Question to ask: Which past project most resembles ours, and what are you assuming about our case?
Domain expertise shortens the ramp-up and helps a firm spot regulatory or operational constraints early. It can also carry assumptions from past clients that hide what is different about yours, so the second half of the question matters as much as the first. A strong answer names specific overlaps with your situation and lists the assumptions the firm plans to test in its first weeks.
Our own projects span financial services and health, public policy and workforce AI adoption, and we find that behavioral patterns often carry over between them while the constraints rarely do.
6. Implementation and product design capability
Question to ask: Who designs and builds what you recommend?
Research only changes behavior once it becomes a screen, a script, a policy or a workflow. A study of 30 US cities in the Journal of Political Economy by Stefano DellaVigna, Woojin Kim and Elizabeth Linos followed 73 randomized trials the cities ran with a national nudge unit. The cities kept the tested nudge in later communications in about 1 in 4 cases (27%), and the strength of the evidence did not strongly predict which ones stuck. The largest predictor was whether the trial changed a communication the city already sent: about two thirds of those were adopted (14 of 21), compared with about 1 in 8 when the trial created something new (6 of 52).
For buyers, a firm's plan for building on your existing systems probably predicts real change better than the size of its pilot results. Ask whether the firm has designers and engineers on staff or hands the build to someone else, and how its recommendations will fit into processes you already run.
Case example: Fortune 500 AI adoption
- Client type: A Fortune 500 HR technology company with a 20,000-person workforce.
- Problem: AI tools had been deployed with training behind them, but usage stayed low.
- Method: A 48-source review of adoption research combined with workforce interviews and surveys. From that, we designed 20 candidate programs and piloted eight with more than 100 employees against a control group.
- Measured result: Confidence in exploring AI tools rose 41% in the pilot group and fell 26% in the control group. Organization-wide AI adoption rose 35%, and use of AI in core responsibilities rose 23%.
- What changed in implementation: The programs ranged from hands-on tasks in real workflows to a challenge series run in the company's existing Slack workspace. Each came with protocols and guides for the client's own teams, and the programs are now scaling to the full workforce.
7. Scalability
Question to ask: What do you expect to happen to the effect when we roll this out to everyone?
A pilot that works with 100 people can fade at 100,000, often because the extra attention of a pilot disappears or because the full population differs from the volunteers who joined first. Ask how the pilot group compares with the full population and what support the intervention needs to keep working without the consultants in the room.
It is worth asking for a forecast as well. In the DellaVigna and Linos study, most forecasters overestimated how well the nudge unit interventions would work, while nudge practitioners were almost perfectly calibrated. A firm that gives you a realistic estimate of the scaled effect before the pilot runs is showing you how it will report results afterward.
Case example: 25-million-member wellness platform
- Client type: A digital wellness platform serving 25 million members.
- Problem: Recommendations ran on a single behavior model applied to every member, whoever they were and wherever they were in their journey.
- Method: A behavioral engine built from scratch that models each member individually and scores every candidate action for that member, tested against synthetic members before it touched the product.
- Measured result: No member outcome figures have been published yet. Pre-launch testing surfaced design failures in an afternoon that would otherwise have appeared months into production.
- What changed in implementation: The engine now powers the platform's recommender in production.
8. Ability to work with internal teams
Question to ask: What will our people be able to run on their own when you leave?
A firm that works alongside your team leaves more behind than one that works in parallel and presents at the end. Ask what your staff will own when the engagement closes, such as protocols, data, code and the skills to design the next experiment, and how the firm will involve your analysts and product owners in decisions along the way.
We design for this from the start. The insurer under question 4 now runs the in-house behavioral science team we designed from interviews with 25 operational stakeholders, and the wellness platform's team owns the engine's specification and adds new actions and member segments without us.
9. Evidence of measurable outcomes
Question to ask: What changed for past clients, and compared with what?
Ask for outcomes stated in full: the behavior measured, the direction of the change, its size, and what it was compared against. Be cautious when a case study reports reach or engagement for a project whose goal was behavior change, and check whether each figure was measured or projected, since both kinds appear in consulting case studies, including ours. In the insurer project, the 12% sales increase was measured against a control group, while the $35 million is a projection built on it.
10. Transparency about uncertainty
Question to ask: What are you unsure about, and what result would change your plan?
Behavioral effects vary with population and context, so any firm that promises a specific result before seeing your data is guessing. Strong firms report the range around an estimate and label clearly which findings come from pre-launch testing and which come from real use. Ask a firm to describe a project where the intervention did not work, and what it changed afterward.
We hold our own case studies to this standard, and the measured-result lines above say plainly where an outcome is still untested.
What to include in a behavioral science RFP
A request for proposals (RFP) that asks the ten questions above in writing makes bids far easier to compare. We suggest including:
- The behavior you want to change, described as a specific action by a specific group, and the business metric it should move.
- What you already know, including past attempts and the data you can share with bidders.
- Constraints the solution must respect, such as regulation, brand guidelines, technical systems and timeline.
- A request for the proposed research and measurement design, with the comparison group and primary outcome named.
- The names and roles of the people who will do the work, and the share of their time allocated to your project.
- Two relevant past experiments with results, including one that failed or came back null.
- An implementation plan that says who designs and builds the intervention and how it fits your existing processes.
- A handover plan listing what your team will own at the end.
- Pricing by phase, with a decision point after each phase, so scale-up is funded only after a pilot shows results.
- The firm's approach to research ethics and data privacy, including how participants give consent.
Warning signs
These patterns do not rule a firm out on their own, but each one deserves a direct follow-up question:
- The proposal quotes effect sizes from famous published studies as if they will apply to your population.
- Every recommended intervention is a nudge, such as a reminder or a default, whatever the underlying problem turns out to be.
- The measurement plan has no comparison group, or defines success as engagement with the intervention and not the behavior it targets.
- Senior experts lead the pitch, but the proposal does not name who will do the work.
- Case studies report numbers without saying what they were compared against.
- The final deliverable is a report, with no plan for building or handing over the work.
- The firm cannot describe an experiment that failed.
- A proprietary framework is presented as validated, with no account of the data it was validated on.
How The Decision Lab approaches these criteria
Our operating model follows the sequence this guide describes. We review the research, test the most promising ideas with the client's own customers or employees, and measure against a control group wherever the setting allows. Our behavioral product design team then builds the winning version into tools the client already uses, and we hand over the protocols and assets so the client's team can keep running it.
The projects above show that model in practice, including our call-center pilot and in-house team design for a Canadian insurer, our AI adoption pilot across a 20,000-person workforce, our behavioral engine for a 25-million-member wellness platform and our controlled framing experiment for a financial wellbeing product. For AI adoption specifically, we use our SPROUT framework, validated with data from more than 20,000 employees, and our AI Adoption Diagnostic applies it to a team through 17 anonymous questions with a minimum of 5 responses.
We are not the right firm for every project. If you need only a short evidence review, a smaller specialist or an academic partner may be a better fit. If you are preparing an RFP for anything larger, we are glad to answer the ten questions above in writing before you shortlist.
Frequently asked questions
How do I choose a behavioral science consulting firm?
Shortlist firms that can show experiments they designed and ran themselves, with results measured against a control group. Then compare how each would turn research into a working intervention and what your own staff could run once the engagement ends. Asking every firm the same ten questions makes their answers comparable.
What should I look for in a behavioral economics consulting firm?
Look for people who read the research critically, since published nudge effects often shrink at scale, and who can test ideas in your own setting before you commit budget to a rollout. Strong firms also bring design and engineering skills, so their recommendations reach customers or employees instead of stopping at a slide deck.
What should I include in a behavioral science RFP?
Describe the behavior you want to change as a specific action by a specific group, and the business metric it should move. Ask each bidder for a measurement design with a named comparison group, the names of the people doing the work, a plan for handing the work to your team, and one past experiment that failed.
How is a behavioral science firm different from a management consultancy?
A behavioral science firm starts from how people actually decide and act, and usually tests changes with experiments before recommending a rollout. A management consultancy typically focuses on strategy and operating structure. Many large projects need both, so ask either kind of firm how it will measure whether behavior changed.
How much does a behavioral science consulting project cost?
Cost depends on how much primary research and experimentation the project needs, and on whether the firm designs and builds the intervention. A literature review with recommendations costs far less than a randomized field trial followed by implementation. Scoping research and build work as separate phases makes competing bids easier to compare line by line.
Which behavioral science consulting firms are best for enterprise work?
For enterprise work, favor firms that have run experiments inside large organizations and then built the winning version into systems those organizations already used. Ask for a reference client of similar size, and ask that client what their team could do on its own once the engagement ended.

