Last updated October 2, 2026
In a 2021 Nature megastudy of exercise programs, researchers led by Katherine Milkman and Angela Duckworth tested 54 programs designed to get 61,293 gym members to exercise more. Just under half raised weekly gym visits above a placebo control group, by 9% to 27%, and forecasters asked in advance could not predict which ones would.
Behavioral science consulting changes behavior an organization can measure, and it is best suited to adoption, conversion, retention and follow-through problems. At The Decision Lab, we are an applied behavioral science firm. We diagnose the barriers behind a behavior, design interventions aimed at them, test those interventions and scale what works, from clean cookstove adoption research for the World Bank to a machine-learning sales pilot for a large insurer that raised sales 12% over a control group. Behavioral science is the wrong first move when the real problem is pricing, access, product quality or strategy.
Controlled testing is the strongest way to show that an intervention caused a result, and the gym study shows why it matters, since ideas that convince experts often do nothing in the field. Not every engagement can or should use a randomized control group, though. Some of our projects rightly end at a diagnosis, a set of customer segments, a prototype or support for a single decision, and the client takes the work forward from there.
Behavioral science consulting, explained
A behavioral science consultancy works out why customers or employees are not doing something an organization needs them to do, then designs and tests changes that make the behavior easier. A company should hire one when people have the information and the incentive to act and still fall short, provided the behavior can be measured and a pilot can be run. A typical project moves from a diagnosis of barriers to designed interventions, then to a pilot measured against a comparison group where possible, and finally to a rollout of what works. The results show up in outcomes such as enrollment, sales, retention and tool adoption, and in better decisions inside the organization. The investment pays off when many people make the decision and each completed behavior is valuable, because a small change across a large group can outweigh the cost of testing.
Behavioral science consulting turns decision research into tested changes
Behavioral science draws on psychology and behavioral economics to explain how people actually make choices. Its central lesson for practitioners is that behavior depends heavily on context. The default option, the timing of a request, the wording of a message, the effort a step demands and what other people appear to be doing all shift what people end up choosing.
The field moved from universities into government first. In a count published in Behavioral Science & Policy, Faisal Naru found 631 organizations applying behavioral science to public policy worldwide in May 2024, up from 201 in 2018, about three times as many. Companies have followed, some by building in-house teams and many by hiring outside firms.
A consultancy's job is to translate published research into changes that work in one particular setting. Research shows which levers have moved behavior elsewhere, and the consultancy works out which of them apply to your customers and employees, then checks the answer against your own data. Our own work spans more than 60 published case studies across health, financial services, education, government and technology, and most of those projects combine behavioral science with a neighboring discipline such as user research or strategy.
Companies should hire behavioral scientists when a measurable behavior falls short
Behavioral science consulting is most useful when an organization has already supplied the information and the incentive, and the behavior still falls short. You should consider it if most of these apply:
- Customers begin a high-value journey, such as onboarding or an application, and do not complete it.
- Employees resist an important tool or process, a pattern we cover in our guide to behavioral barriers to enterprise AI adoption.
- Retention, repayment, enrollment or follow-through is running below target.
- You can name the target behavior and count how often it happens.
- You can run a pilot, ideally with a comparison group that does not receive the change.
- Improving the behavior is worth more than the cost of testing changes to it, a calculation we set out in the section on whether behavioral science consulting is worth it.
Two further kinds of problem suit the approach. Committees and investment teams make decisions that bias or deadlock pulls off course, and digital health and personal finance products deliver value only if people keep using them. In every case the gap sits at the moment of decision, where a small change in design, such as a different default or a well-timed reminder, can shift what large numbers of people do.
Behavioral science is the wrong first step when the barrier is structural
Nick Chater and George Loewenstein, two researchers who had long subscribed to the approach themselves, argued in Behavioral and Brain Sciences that behavioral public policy leaned too heavily on individual-level fixes for problems that needed changes to the system people operate in. Companies face the same risk: if customers cannot afford a product, or the product does not do its job, better messaging will not close the gap, and a behavioral project can delay the decision that would.
Behavioral science is usually the wrong first step in four situations:
- The barrier is price or access. Pricing or distribution decisions come first.
- The product or service does not work well. Product development and usability research come first.
- The open question is what to do, such as which market to enter or which product to build. Management consulting is built for strategy questions of this kind, and we compare the disciplines in our guide to behavioral science vs. UX research, market research and management consulting.
- Nobody can measure the behavior. Without a way to count it, no one can tell whether a change worked, so the first step is building that measurement.
A good diagnosis often finds structural and behavioral barriers side by side. When we mapped barriers to clean cookstove adoption in Uganda for the World Bank, we identified 26 of them, ranging from low awareness of the health risks of indoor smoke to the weight of the stoves themselves. The cheapest model was too heavy for most women to carry on their own. Messaging could not have solved the weight problem, and the diagnosis made that clear before any money went into a campaign.
Methods range from interviews and surveys to controlled experiments and simulation
Most projects combine several methods, chosen for what the organization needs to know. These are the ones we use most:
- Evidence reviews. A structured read of published research on the behavior, used to generate hypotheses and to avoid reinventing ideas that have already been tested. For a Fortune 500 AI adoption project, our review covered 48 sources.
- Interviews and field observation. Conversations with the people whose behavior matters and the staff who serve them, plus direct observation of where decisions happen.
- Behavioral diagnostics. A structured map of the barriers and drivers behind a behavior, often organized with an established framework such as the COM-B model of behavior change.
- Analysis of existing behavioral data. App events and call recordings show what people do without asking them. For one of North America's largest insurers, we analyzed more than 170,000 hours of sales calls.
- Surveys and segmentation. Surveys measure attitudes and self-reported habits at scale, and machine learning can then find groups of people that demographics miss.
- Controlled experiments. Randomized controlled trials and A/B tests assign people at random to receive a change or not, so any difference in behavior can be traced to the change itself. When random assignment is not possible, a matched comparison group or a careful before-and-after measure is the next best evidence, with its limits stated.
- Choice experiments. A discrete choice experiment asks people to pick between options that vary on several features at once, revealing which trade-offs they will accept.
- Simulation. Testing a design against synthetic users before launch catches failures early. For a digital wellness platform serving 25 million members, we ran a full behavioral recommendation engine against simulated members before it reached the product.
Qualitative research explains why a behavior happens, and experiments show whether a change moves it. A project that skips the first tends to test the wrong ideas, and one that skips the second ends with a recommendation nobody has checked.
Most projects move from diagnosis to design to testing to scale
A full engagement usually runs in four stages. Our work on AI adoption across a 20,000-person Fortune 500 workforce, for an HR technology company, went through all four.
Diagnose
The team defines the target behavior precisely and maps how people move toward it, from first contact to habitual use. It then identifies the barriers at each step and which groups face which barriers. At the HR technology company, the diagnosis found five barriers and two segments with opposite problems: enthusiasts who were already experimenting and hit friction, and skeptics who distrusted AI and stalled at first contact.
Design
The team generates candidate changes, each aimed at a specific barrier for a specific group, and ranks them on evidence and feasibility. We designed 20 candidate programs for the HR technology company and selected eight for piloting, ranging from hands-on tasks built into real workflows to a mission-based challenge series run in Slack.
Test
The strongest candidates run as pilots against a control group, with outcomes measured before and after. The HR technology company's eight programs ran for one month with more than 100 employees. Confidence in exploring AI tools, measured before and after the pilot, rose 41% among participants and fell 26% in the control group over the same period.
Scale
Changes that work are rolled out with the protocols and guides the client's teams need to run them without us. The HR technology company's programs are now scaling to its full workforce.
Many engagements cover only part of this sequence. Some end with a diagnosis the client acts on internally, and others begin at the testing stage because the client already has candidate ideas and wants to know which one works.
Typical outputs include diagnostics, intervention designs, prototypes and experiment results
What a client receives depends on where the project stops, but most deliverables fall into six types:
- A behavioral diagnostic. The barriers and drivers behind the behavior, ranked by how much each matters and how easily it can be changed.
- Segments or behavioral profiles. Groups defined by what drives their behavior. For the Canadian Cancer Society, our analysis showed that whether smoking is part of someone's identity separates smokers more sharply than their tolerance for risk or the support around them.
- Frameworks. Reusable structures for future decisions. For EdReports, we built an evidence uptake framework for school district purchasing around five drivers of whether purchasers use evidence when they choose instructional materials.
- Intervention designs and prototypes. Scripts, messages, product features, wireframes or high-fidelity mockups, ready for testing.
- Experiment results. The size of each change's effect against a control group, including the changes that did not work.
- Working systems and new capability. A production system or an in-house team. For one of North America's largest insurers, we designed a permanent behavioral science team, down to its hiring profiles, alongside the first project.
Whatever the format, the client should own it when the engagement ends. Our digital wellness client received the full specification of its behavioral engine along with the engine itself, and its own team now adds new actions and member segments without a rebuild.
Good results are measured in behavior, against a control group
A strong result has five features, and a buyer can ask to see each of them.
- It measures a behavior the organization can count, such as purchases, sign-ups, visits or tool use. Surveys of intention help with diagnosis, but they make a weak final outcome.
- It compares against a control group wherever one is possible. Outcomes drift for reasons unrelated to any intervention, and the comparison removes that drift. In the HR technology pilot, confidence fell in the control group over the same month, so looking only at participants before and after would have understated the effect.
- It is sized for the real world. Stefano DellaVigna and Elizabeth Linos compared nudge trials published in academic journals with 126 trials run at scale by two US government nudge units. The published trials raised take-up, meaning the share of people who did the target behavior, by 8.7 percentage points over their control groups on average. The at-scale trials raised it by 1.4 points, about a sixth as much.
- It lasts. In the gym megastudy, only about 1 in 12 programs still produced a measurable change in gym visits once the four weeks were over, compared with just under half while the programs were running.
- It connects to a business outcome. The insurer's sales pilot translated its 12% rise over the control group into a projected $35 million increase in annual revenue.
We also report the changes that did not work. A null result tells a client where not to spend money, and a consultancy that only shows its successes is probably showing a small part of its record.
Whether behavioral science consulting is worth it depends on how many people the behavior touches
The return on a behavioral project can be estimated before it starts:
Estimated value = decisions affected × realistic behavior change × value per completed behavior − implementation cost
Take a hypothetical subscription business in which 200,000 customers reach a renewal decision each year and each renewal is worth $300 in annual revenue. Suppose a tested change to the renewal flow raises the renewal rate by 1.5 percentage points over its current level, close to the 1.4-point average DellaVigna and Linos found for nudges run at scale. The business would keep 3,000 more customers a year, worth about $900,000, against a hypothetical project and rollout cost of $250,000.
The same change at a business with 20,000 customers would bring in about $90,000 and would not cover that cost. Between organizations, the number of decisions affected usually varies far more than the realistic size of the effect, so it tends to settle the case. Many behavioral changes, such as a revised call script or a different default, also cost little to roll out once they are proven. The insurer's intervention was a set of script changes and agent training applied to calls the company was already making.
Projects are harder to justify when few people make the decision, or when the organization cannot test changes and measure the result. In those cases a diagnosis on its own can still pay for itself by stopping spending on a solution aimed at the wrong barrier.
Our case studies show the method at work in different sectors
For one of North America's largest insurers, we used process mapping to locate the critical moments in agent sales calls and machine learning to cluster where and how those calls went wrong. We then built script changes and agent training around the clusters and piloted them against a control group. The case study on recovering lost insurance sales with machine learning describes the full design, including the in-house team we set up alongside it.
For the Canadian Cancer Society's text-message quit-support service, we surveyed 122 smokers and ex-smokers and used decision trees and cluster analysis to find four distinct groups. Each group received its own conversation flow, with message timing tied to the moments of the day each user had flagged as hardest. Our write-up on helping smokers quit with machine learning explains how the profiles carry over to the charity's other programs.
Health Canada asked us to help the steering committee for organ donation and transplantation agree on a national governance framework. We used interviews and a discrete choice experiment to rank the options, then ran a workshop in which members reallocated points across solutions until the plan reflected the range of views on the committee. The case study on supporting organ donation governance across Canada covers both phases.
For Money Guided, an employee financial wellbeing product, we surveyed 533 UK employees and 264 HR managers and ran a 400-person experiment comparing four ways of framing the product's value against a control. The winning frame now shapes how the product presents itself to employees and to the HR managers who buy it. We describe this result in our case study on reducing stress-driven money mistakes as design validation measured before launch, with in-product behavior as the next test.
For a leading online learning platform, we surveyed more than 800 students and used unsupervised machine learning to group them. Willingness to subscribe, an attitude, produced cleaner and more predictive segments than any demographic, and the resulting six segments now inform the platform's product strategy, as our case study on keeping students learning in the age of AI describes.
A good consultancy can show its experiments and hand over what it builds
When you compare firms, four questions separate them quickly:
- Can the firm show experiments it ran with control groups, including ones that did not work?
- Will it test ideas with your own customers or employees, or rely on published findings?
- Can it design and build the change, or only recommend it?
- What will your team own when the engagement ends?
Our guide on how to choose a behavioral science consulting firm covers ten questions in more depth, and they are the questions we would want any buyer to ask us as well as every other firm on a shortlist.
Before issuing a brief or a request for proposals, it helps to write down four things: the behavior you want to change, the number you will use to measure it, the data the firm will be able to access, and who in your organization can approve a test. In our experience, briefs that include these produce proposals that are easier to compare, and they shorten the diagnosis stage of whatever project follows.
Frequently asked questions about behavioral science consulting
What does a behavioral science consultancy do?
A behavioral science consultancy helps an organization change a specific behavior, such as enrolling in a plan or renewing a contract, by finding the barriers behind it and testing changes against a control group. Most firms also design the product and message changes that follow, and some build the software or in-house teams that carry the work forward.
When should a company hire behavioral scientists?
Hire behavioral scientists when a behavior that matters to the business falls short even though people have the information and the incentive to act. The clearest signal is a gap your data already shows, such as a sign-up flow that loses half its users at one step, combined with the ability to test changes on real customers or employees.
How can behavioral science help a company?
Behavioral science helps a company in two broad areas. With customers, it improves outcomes such as sales and retention. Inside the organization, it improves tool adoption and the quality of decisions made by committees and teams. The internal side gets less attention, although employee behavior shapes the return on many internal investments, including AI.
What does a behavioral science project look like?
Most projects pair behavioral scientists with designers or data scientists, working alongside a client sponsor who can grant access to data and approve tests. The work moves from a diagnosis to designed changes and then to a pilot against a control group, and the client decides at the end of each stage whether to continue to the next one.
Is behavioral science consulting worth it?
It is worth it when many people make the decision and each decision carries real value, since small changes add up across a large base. It is harder to justify for rare behaviors or ones nobody can measure. Before signing, ask the firm to estimate what a realistic size of change would be worth to you, and check its assumptions.
What is the difference between behavioral science and behavioral economics?
Behavioral economics is one branch of behavioral science. It studies how real decisions depart from the predictions of standard economic models, for example through loss aversion, the tendency for a loss to weigh more heavily than a gain of the same size. Consultancies draw on it alongside psychology and data science, and in practice the two terms are used interchangeably.
Can a company build its own behavioral science team instead?
A company can build its own team, and a common path is to start with an outside firm on one high-value project, then use what it proves to justify hiring. We followed that route with one of North America's largest insurers, designing the structure and hiring profiles of its permanent team while running the first pilot that tested the team's value.

