Does Behavioral Science Work in Business? Evidence, ROI, and Real Examples

Last updated October 2, 2026

In a 2021 field experiment published in Nature, Katherine Milkman and colleagues compared 54 four-week programs designed to get 61,293 members of a US gym chain to exercise more. Just under half of the programs raised weekly gym visits over a control group, by between 9% and 27%, and impartial forecasters could not predict in advance which programs would be the ones that worked (Milkman et al., 2021).

Behavioral science works in business when it changes how a decision is structured at a moment that already drives revenue or cost, and when that change is tested against a control group and counted in business units such as sales or retained customers. At The Decision Lab, we are an applied behavioral science firm, and our clearest commercial result came from that combination: redesigned sales calls at a top Canadian insurer raised sales 12% over a control group, with a projected $35 million increase in annual revenue. The evidence is weaker for messages that only give people information, and for effects measured long after an intervention ends.

In this guide, behavioral science includes behavioral economics, choice architecture, nudging and related methods for changing real-world decisions. This page sets the published research and large external field experiments alongside our own case studies, each labeled as a TDL case. For every result we state the business outcome, the intervention, the sample and context, the measured effect and the limits of the evidence.

The short answer: behavioral science works in business under four conditions

Behavioral science works in business when an intervention changes how a decision is structured at a moment that already drives revenue or cost, and its effect is measured against a control group. Our clearest example is a 12% sales lift for a top Canadian insurer. At scale, expect effects about a sixth the size journals report.

  1. The intervention changes the structure of a choice, such as a default, a form or a sales script, instead of only adding information.
  2. It targets a moment that already drives a business outcome, such as enrollment, renewal, a sales conversation or onboarding.
  3. It is tested against a control group in the organization's own setting before anyone counts the return.
  4. The return is counted in business units and compared with the full cost of the intervention.

When one of these is missing, both the published evidence and our own projects show smaller or less certain returns.

How to calculate the ROI of behavioral science

We calculate the return on any behavioral intervention the same way: the value of the behavior change a control group shows the intervention caused, minus its full cost, divided by that cost.

ROI = ((measured lift × decisions affected × value per decision × share that persists) - full cost) ÷ full cost

Measured lift is the difference between the treated group and the control group, such as a 2-point rise in renewal rate. Value per decision is what one more sale or retained client is worth, and the share that persists accounts for effects that fade once the intervention stops. Full cost covers staff time, vendor fees, engineering work and compliance review.

As a hypothetical illustration, a 2-point rise in renewals across 100,000 customers worth $500 each, with half the effect persisting, is worth $500,000. Against a full cost of $100,000, that is a return of $4 for every $1 spent, on top of recovering the cost.

Examples of behavioral science producing business results

  1. Retirement plan enrollment: automatic enrollment in a large US company's 401(k) plan raised participation among new employees from 37% under opt-in to 86%.
  2. Household energy use: Opower's neighbor-comparison reports cut electricity use by 2.0% against control households, across about 600,000 homes.
  3. Overdue payments: one added sentence in UK tax reminder letters, saying most people pay on time, raised payment rates within 23 days by between 1.3% and 2.1%.
  4. Sales revenue, TDL case: our sales-call redesign for a top Canadian insurer raised sales 12% over a control group in the first pilot.
  5. Client retention, TDL case: our dropout prevention work for American Financial Solutions cut the dropout rate from its debt management program by half.

Behavioral science can increase revenue when it changes the sales conversation itself

None of the external field experiments on this page measures company revenue directly, so our insurer project is the clearest revenue evidence we can offer. One of North America's largest insurers asked us to design a permanent in-house behavioral science team and prove its value on a first project.

We analyzed more than 170,000 hours of agent sales calls, mapping the critical moments in each conversation and using machine learning to group the points where calls went wrong. From those groups we built targeted script changes and agent training, then piloted them against a control group. Sales rose 12% over the control group in the first pilot, and the insurer sales-call case study reports a projected $35 million increase in annual revenue. The insurer now runs its own behavioral science team, which we designed from interviews with 25 operational stakeholders.

The $35 million is a projection from one pilot, and the published case study does not state the pilot's size or duration. A buyer using a result like this as a benchmark should ask how long the lift held after the pilot ended, and whether it was measured across every agent group or only the pilot teams.

The research evidence on behavioral economics ROI is real and uneven

The largest review of nudges, published in the Proceedings of the National Academy of Sciences by Stephanie Mertens and colleagues in 2022, pooled more than 200 studies covering about 2.1 million people (Mertens et al., 2022). The average intervention shifted behavior by just under half a standard deviation compared with a control group, a size the authors describe as small to medium. Food choices responded most, with effects up to 2.5 times larger than in other areas, and the authors found that positive results were more likely to reach publication than null ones.

Maximilian Maier and colleagues reanalyzed the same studies later that year, correcting for the null results missing from the literature, and found no remaining evidence that the average nudge changes behavior (Maier et al., 2022). They found evidence against any effect for interventions that only give people information or help them follow through, while the evidence for interventions that change the structure of the choice itself, such as defaults, was inconclusive.

Defaults have the strongest record in this literature. A meta-analysis of 58 default studies with 73,675 participants in total, by Jon Jachimowicz, Shannon Duncan, Elke Weber and Eric Johnson, found that pre-selecting an option moved choices toward it by about two-thirds of a standard deviation on average (Jachimowicz et al., 2019). Results varied widely across those studies, and two of them moved choices in the wrong direction.

The gap between published studies and real deployments is the number a business should plan around. Stefano DellaVigna and Elizabeth Linos compared 74 nudges published in academic journals with 126 randomized trials run at scale by two US government nudge units, covering 23 million people (DellaVigna and Linos, 2022). Published nudges raised take-up, meaning the share of people who did the target behavior, by an average of 8.7 percentage points over a control group. The at-scale trials raised it by 1.4 points, about a sixth as much.

We treat that at-scale figure as the realistic expectation for a client's first test. A lift of a point or two can still pay back many times over when the intervention costs little and the decision is made millions of times, and the field experiments below show that arithmetic at work.

External field experiments show where behavioral science produces business results

Automatic enrollment more than doubled retirement plan participation

When a large US company switched its 401(k) retirement plan from opt-in to automatic enrollment, Brigitte Madrian and Dennis Shea found that participation among newly eligible employees rose from 37% under opt-in to 86% under automatic enrollment (Madrian and Shea, 2001). Many employees then stayed at the low default contribution rate and the default fund the company had chosen, so participation rose further than savings did.

Neighbor comparisons cut household electricity use by about 2%

Opower, a company that sends utility customers reports comparing their electricity use with their neighbors', ran randomized trials across about 600,000 treatment and control households in the US. Hunt Allcott found that the reports cut electricity use by 2.0% on average against control households, an effect equal to a short-run price rise of 11% to 20% (Allcott, 2011). The heaviest users cut their use by 6.3% and the lightest by 0.3%. A follow-up by Allcott and Todd Rogers found that savings faded between reports but continued after the reports stopped (Allcott and Rogers, 2014).

One added sentence in a reminder letter sped up tax payments

Working with the UK tax authority, Michael Hallsworth, John List, Robert Metcalfe and Ivo Vlaev tested reminder letters on more than 200,000 people who owed overdue tax (Hallsworth et al., 2017). In the first trial, adding a sentence saying that most people pay on time raised payment rates within 23 days by between 1.3% and 2.1% over the standard letter, depending on the wording, at almost no added cost. Statements about what other people do outperformed statements about what people ought to do. We would expect the same mechanics to apply to a company chasing overdue invoices or lapsed renewals, though that transfer has to be tested.

Behavioral tools often beat incentives per dollar, once costs are counted carefully

Shlomo Benartzi and colleagues compared behavioral interventions with financial incentives and education, measuring impact per dollar spent (Benartzi et al., 2017). In retirement savings, an enrollment form asking new employees to make an active choice generated about $100 in additional savings per dollar spent, against $1.24 per dollar for tax credits. Avishalom Tor and Jonathan Klick later argued that the analysis counted costs inconsistently in ways that favored nudges, and that organizations should run full cost-benefit comparisons instead of assuming behavioral tools win (Tor and Klick, 2022).

About 1 in 12 gym programs kept working after they ended

In the Milkman megastudy, only 8% of the 54 programs produced a change in gym visits that was still measurable after the four weeks ended. The top performer gave small rewards for returning to the gym after a missed workout.

TDL case studies: measured business outcomes and the strength of each result

How we selected and graded this evidence

From the 63 case studies on our behavioral science case studies page, we selected six that report a quantified outcome tied to a business result and that together cover the main kinds of evidence we produce. Each figure is stated as the case study reports it. We tag every result with its evidence type and treat a controlled comparison in the client's live setting as the strongest evidence of return, since before-and-after comparisons and pre-launch experiments cannot show what would have happened without the change.

The evidence types used below are Controlled comparison, Before-and-after, Simulated environment, Self-reported outcome, Pre-launch experiment and Comparison not described.

TDL case studies by business outcome and evidence type
TDL case Business outcome Intervention Sample and context Measured effect Evidence type Main limitation
Insurer sales-call redesign Sales revenue Script changes and agent training, based on analysis of 170,000+ call hours Agent sales calls at a top Canadian insurer Sales up 12% vs control; projected $35M more revenue a year Controlled comparison Revenue projected from one pilot; pilot size not published
American Financial Solutions dropout prevention Client retention Risk model plus redesigned counselor scripts and client messages Nonprofit debt program; records from thousands of clients Dropout rate down 50% Comparison not described No control group or time window published
Fortune 500 AI adoption pilot Employee AI use Eight programs, each aimed at one adoption barrier 100+ employees for one month in a 20,000-person company Confidence up 41% vs down 26% in control Controlled comparison Self-reported outcome Confidence is self-reported over one month
Consumer electronics design process Design efficiency Simulation-first behavioral design process New feature in a health app with 60 million users Design time down 75% Before-and-after No comparison team; consumer behavior not yet measured
Pinterest shopper decision study Evidence for advertiser sales Five experiments in working platform replicas 2,400 simulated shopping sessions 29 quantified Pinterest strengths Simulated environment No revenue effect published
Money Guided product framing Product positioning Four value framings tested against a control 400-person controlled experiment Winning frame shaped product messaging Pre-launch experiment In-product behavior not yet tested

The insurer pilot is the only result here that reports a revenue effect against a control group. The others measure outcomes that sit upstream of revenue, such as retention or design speed, and converting them into a dollar figure takes assumptions a buyer should ask any firm, including us, to state.

Behavioral science produces the clearest ROI at high-volume decisions with a structural fix

The strongest returns in the research and in our projects come from changing what a person is presented with at the moment of choice. Automatic enrollment and rewritten sales scripts change the path of least resistance, while a reminder or an information leaflet asks people to do more work. The Maier reanalysis is consistent with this pattern, since information-only interventions were the category with evidence against an effect.

Volume turns small effects into large sums. A 2% cut in electricity use is modest for one household and substantial across 600,000 of them, and the tax letters moved payment rates by a couple of percent across more than 200,000 people at almost no added cost. The projected revenue from our insurer pilot rests on the same arithmetic.

The moment matters as much as the mechanism. In our dropout prevention work for American Financial Solutions, the model traced later dropout back to the initial onboarding call, so the redesign started there instead of at the point where clients left. Returns are clearest when the intervention sits at a decision that already has a known value per unit, such as a sale or a retained client, because the effect can then be priced without heroic assumptions.

Returns also last longer when the client keeps the capability. The insurer now runs its own behavioral science team, and our consumer electronics client owns the simulation-first design process it used to cut design time by three-quarters. A one-off intervention has to pay back on its own, while a team that keeps testing can compound small gains across many decisions.

The evidence is weaker for information-only messages and for lasting effects

Messages that tell people facts or remind them of a goal are the cheapest interventions to run and the ones with the least support once publication bias is corrected. We still use them, usually as one variant in a test alongside a structural change, and we expect them to lose more often than they win.

Persistence is the largest open question for business cases that assume a permanent lift. Only about 1 in 12 programs in the gym megastudy produced a change still visible after it ended, and Opower's savings dipped between reports. A business case built on year-one effects carrying into year three needs a measurement plan that checks whether they do.

Transfer from one setting to another is weaker than headline studies imply. The sixfold gap between journal nudges and at-scale government trials probably reflects both which results get published and the difference between a carefully run study and a program run by an organization alongside its other work. Averages also hide wide variation: Opower's heaviest users responded about 20 times as strongly as its lightest, and some default studies moved choices the wrong way.

Some of our own evidence sits at the weaker end too. Self-reported confidence and pre-launch framing tests are useful for deciding what to build, and none of them is a measured return. We report them as what they are, and we push clients to follow them with an in-market test.

How to measure the ROI of behavioral science in seven steps

  1. Name the business outcome and its value per unit before designing anything, for example revenue per completed sale or lifetime value per retained client.
  2. Write down the expected effect using at-scale benchmarks, such as the 1.4-point average from government nudge units, and check whether the business case still works at that size.
  3. Design several candidate interventions, since forecasters in the megastudy could not tell in advance which programs would work.
  4. Hold out a randomized control group in the live setting, or a matched comparison group when randomization is impossible, and fix the measurement window in advance.
  5. Measure behavior from system records, such as sales or payment data, and treat survey and self-report measures as supporting evidence.
  6. Keep measuring after the intervention ends to see how much of the effect persists, and price only the part that does.
  7. Calculate return as the value of the measured lift minus the full cost, including staff time, vendor fees, engineering work and compliance review, and compare it per dollar with the option the organization would otherwise fund.

We also recommend reporting interventions that showed no effect. An organization that only records its wins will overestimate the return on its next project in the same way the published literature overestimates the average nudge.

How we measure impact in our behavioral science projects

We start each engagement by agreeing with the client on the outcome that counts and how it will be measured, before any intervention is designed. For the insurer, we ranked candidate pilots on criteria including expected return before choosing agent sales calls as the first test, and we ran that test against a control group.

We design more interventions than we expect to keep. In our Fortune 500 AI adoption work, we built 20 candidate programs, each aimed at one barrier for one group of employees, and piloted eight of them against a control group. When a project can only validate a design before launch, as our Money Guided product framing work did, we say so in the results and name the in-market test that should follow.

We also share the at-scale benchmark with clients at the outset, so the first test is judged against a realistic expectation of a point or two of added take-up instead of the larger effects common in journals. Our aim is for clients to keep the method once we leave, and the insurer's in-house behavioral science team is the clearest example of that.

Frequently asked questions

Does behavioral science work in business?

Behavioral science works in business when an intervention changes how a decision is structured and is tested against a control group. Field experiments show measurable effects on retirement plan participation and household energy use. Effects at scale average about a sixth of those in journal studies, so we advise testing several variants and expecting some to show nothing.

What is the ROI of behavioral economics?

There is no single figure, because the return depends on the value of the decision being changed and the cost of changing it. One widely cited analysis found an active-choice enrollment form produced about $100 in added retirement savings per dollar spent, though critics argue such ratios undercount costs. A credible estimate also prices how long the effect lasts.

Can behavioral science increase revenue?

It can, especially where revenue depends on a repeated decision such as a sales conversation or a renewal. In our work with a top Canadian insurer, script changes and agent training raised sales 12% over a control group in the first pilot. A projection such as the $35 million annual figure depends on the effect holding at full rollout.

What are examples of behavioral science producing business results?

Documented examples include automatic enrollment raising retirement plan participation among new employees from 37% to 86%, and neighbor-comparison reports cutting household electricity use by 2%. TDL cases include a 12% sales lift for an insurer. When judging any example, check that it reports a comparison group, because a before-and-after change can reflect seasonality or unrelated changes.

How long do behavioral science effects last?

Effects often fade without repetition: in the 54-program gym megastudy, about 1 in 12 programs produced a change still measurable after the four weeks ended, and Opower's energy savings dipped between reports. We expect structural changes such as defaults to last longer, because they keep working without anyone having to remember them, but that still needs measuring.

How large an effect should a company expect from a nudge?

Plan around at-scale results rather than journal averages: across 126 government nudge-unit trials, the average rise in take-up was 1.4 percentage points over control groups. In high-volume settings that can be worth a great deal. A business case that only works at an 8- or 9-point lift depends on effects that journals report far more often than real deployments deliver.

Sources

  1. Mertens, S., Herberz, M., Hahnel, U. J. J., & Brosch, T. (2022). The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains. Proceedings of the National Academy of Sciences, 119(1).
  2. Maier, M., Bartoš, F., Stanley, T. D., Shanks, D. R., Harris, A. J. L., & Wagenmakers, E.-J. (2022). No evidence for nudging after adjusting for publication bias. Proceedings of the National Academy of Sciences, 119(31).
  3. Jachimowicz, J. M., Duncan, S., Weber, E. U., & Johnson, E. J. (2019). When and why defaults influence decisions: A meta-analysis of default effects. Behavioural Public Policy, 3(2), 159-186.
  4. DellaVigna, S., & Linos, E. (2022). RCTs to Scale: Comprehensive Evidence from Two Nudge Units. Econometrica, 90(1), 81-116.
  5. Benartzi, S., et al. (2017). Should governments invest more in nudging? Psychological Science, 28(8), 1041-1055.
  6. Tor, A., & Klick, J. (2022). When should governments invest more in nudging? Revisiting Benartzi et al. (2017). Review of Law and Economics, 18(3), 347-376.
  7. Milkman, K. L., et al. (2021). Megastudies improve the impact of applied behavioural science. Nature, 600, 478-483.
  8. Madrian, B. C., & Shea, D. F. (2001). The power of suggestion: Inertia in 401(k) participation and savings behavior. Quarterly Journal of Economics, 116(4), 1149-1187.
  9. Allcott, H. (2011). Social norms and energy conservation. Journal of Public Economics, 1082-1095.
  10. Allcott, H., & Rogers, T. (2014). The short-run and long-run effects of behavioral interventions: Experimental evidence from energy conservation. American Economic Review.
  11. Hallsworth, M., List, J. A., Metcalfe, R. D., & Vlaev, I. (2017). The behavioralist as tax collector: Using natural field experiments to enhance tax compliance. Journal of Public Economics, 148, 14-31.
  12. The Decision Lab. Recovering $35 million in lost sales using machine learning (case study).
  13. The Decision Lab. Preventing debt repayment dropouts using predictive modeling (case study).
  14. The Decision Lab. Increasing AI adoption across a 20,000-person workforce (case study).
  15. The Decision Lab. Reducing product feature design time using simulation-first validation (case study).
  16. The Decision Lab. Understanding how shoppers decide on Pinterest by simulating the entire purchase journey (case study).
  17. The Decision Lab. Reducing stress-driven money mistakes using large-scale consumer experiments (case study).

About the Authors

A man in a blue, striped shirt smiles while standing indoors, surrounded by green plants and modern office decor.

Dan Pilat

Managing Director

Dan is a Co-Founder and Managing Director at The Decision Lab. He is a bestselling author of Intention - a book he wrote with Wiley on the mindful application of behavioral science in organizations. Dan has a background in organizational decision making, with a BComm in Decision & Information Systems from McGill University. He has worked on enterprise-level behavioral architecture at TD Securities and BMO Capital Markets, where he advised management on the implementation of systems processing billions of dollars per week. Driven by an appetite for the latest in technology, Dan created a course on business intelligence and lectured at McGill University, and has applied behavioral science to topics such as augmented and virtual reality.

A smiling man stands in an office, wearing a dark blazer and black shirt, with plants and glass-walled rooms in the background.

Dr. Sekoul Krastev

Managing Director & Co-Founder

Dr. Sekoul Krastev is a decision scientist and Co-Founder of The Decision Lab, one of the world's leading behavioral science consultancies. His team works with large organizations—Fortune 500 companies, governments, foundations and supernationals—to apply behavioral science and decision theory for social good. He holds a PhD in neuroscience from McGill University and is currently a visiting scholar at NYU. His work has been featured in academic journals as well as in The New York Times, Forbes, and Bloomberg. He is also the author of Intention (Wiley, 2024), a bestselling book on the science of human agency. Before founding The Decision Lab, he worked at the Boston Consulting Group and Google.

Notes illustration

Eager to learn about how behavioral science can help your organization?