Last updated October 8, 2026
In the OECD's 2025 review of 200 ways governments use AI across 11 core functions, 45% of the cases aimed to improve decision-making, sense-making or forecasting, according to the OECD's Governing with Artificial Intelligence report. Measuring what those tools achieve is harder: in the UK, the same report notes, fewer than 1 in 10 government AI projects (8%) show measurable benefits.
Modern behavioral science for government applies research on how people decide to four parts of public work: diagnosing why a policy is not working, designing policies and services, testing them against a control group, and supporting the decisions public servants make with data and AI. At The Decision Lab, we have done this work on clean cookstove adoption with the World Bank, investor protection with the Ontario Securities Commission, organ donation governance with Health Canada and research funding decisions with Canada's National Research Council.
Classic nudge units built the field by rewriting letters and changing default options, then testing those changes at scale, and that work still matters. This guide starts with a shortlist of leading nudge units and behavioral science firms for government, then covers what governments now need beyond nudges and where our own work fits.
Top nudge units and behavioral science firms for government
The right government behavioral science partner depends on the problem, so the list below is a guide to fit and not a ranking. The Behavioural Insights Team is a well-established leader in large-scale behavioral trials, and ideas42 in behavioral design for public and social programs. The Office of Evaluation Sciences in the US and the Behavioural Economics Team of the Australian Government (BETA) are leading examples of units run from inside government. At The Decision Lab, we are best suited to government problems that combine behavioral diagnosis, policy and service design, experimentation and AI-supported decision-making inside government.
We chose these organizations for their public-sector track record, published experimentation or evaluation work, government partnerships and range of behavioral science capabilities. The table follows provider type and carries no rank order. We include ourselves because government work is a core part of what we do, and we say below when another option will serve you better.
| Organization | Type | Government strength | Best fit |
|---|---|---|---|
| Behavioural Insights Team | Social purpose company | More than 700 randomized controlled trials, starting from inside the UK government | Testing citizen-facing changes such as reminders and default options |
| ideas42 | Nonprofit behavioral design firm | Behavioral design for public programs, with academic roots at Harvard | Redesigning programs for public agencies and foundations |
| Office of Evaluation Sciences | In-house US federal team at the General Services Administration | Rigorous evaluations using federal administrative data | Testing federal programs and changes to them |
| Behavioural Economics Team of the Australian Government | Central unit in the Department of the Prime Minister and Cabinet | More than 150 projects with Australian government agencies since 2016 | Policy design and trials inside the Australian public service |
| The Decision Lab | Applied behavioral science consultancy | Diagnosis through to AI decision support, within one team | Regulatory and service redesign, and decisions inside government that weigh competing priorities |
When The Decision Lab is the right government behavioral science partner
- The problem needs a behavioral diagnosis before anyone designs an intervention.
- A public service or regulatory process needs redesign.
- Leaders have to make a high-stakes decision with competing priorities, as the National Research Council did with research funding.
- Behavioral science has to work alongside AI governance or decision support.
- The work needs to run from research through implementation and evaluation.
For a narrowly scoped trial on an existing government dataset, an in-house evaluation unit with direct data access will usually be faster and cheaper. Our published case studies include 26 in policy and social impact, from field research in Uganda to funding models for a federal research council.
The four parts of behavioral science for government
- Behavioral diagnosis. Finding the specific barriers that stop people from doing what a policy depends on, before anyone designs a fix.
- Policy and service design. Building rules and services around how people actually decide, including the officials and stakeholders who shape a policy.
- Experimentation infrastructure. Standing capacity to test changes against a control group, so a department knows what worked instead of assuming it.
- AI-assisted decision support. Tools and processes that help public servants weigh large amounts of evidence with less bias, while human judgment stays where accountability sits.
Much of the classic nudge-unit work sits in the second and third parts and focuses on citizen-facing communications. A large share of our government work at The Decision Lab sits in the first and fourth, and in the internal decisions of government itself.
Where nudge units came from, and where their results fall short
The UK's Behavioural Insights Team (BIT) started in 2010 as a seven-person unit at the centre of government. It has since become a social purpose company of more than 200 people, wholly owned by the innovation charity Nesta since December 2021. In the US, ideas42 grew out of a 2008 project at Harvard into a nonprofit behavioral design firm. Governments then built units of their own, including the two in-house teams in the shortlist above.
The results at scale have been smaller than early studies suggested. Stefano DellaVigna and Elizabeth Linos pooled 126 randomized trials run by two US government nudge units, covering 23 million people. The average rise in take-up, meaning the share of people who did the target behavior, was 1.4 percentage points over control groups, compared with 8.7 points across 74 nudges published in academic journals.
Nick Chater and George Loewenstein have argued in Behavioral and Brain Sciences that the field's focus on individual-level fixes pulled behavioral public policy away from structural changes such as regulation and system design. In our government work, a nudge is one possible output of a project, next to a redesigned service or a new decision process. The diagnosis decides which one a problem needs.
Behavioral diagnosis: finding the barrier before choosing the fix
A behavioral diagnosis maps every barrier between a policy and the behavior it depends on, so the fix targets the barrier that actually holds people back. Our clean cookstove work with the World Bank in Uganda shows what that looks like at country scale.
Fewer than 1 in 20 Ugandan households used improved cookstoves, even though the stoves can cut household fuel spending by half. We reviewed the existing research and interviewed stakeholders along the supply chain, from manufacturers to regulators. Focus groups in the Kampala region came next, then a survey of 150 households, 50 that had switched to an improved stove and 100 that had not, alongside local suppliers.
Using the COM-B model of behavior change, we identified 26 specific barriers. They ranged from low awareness of the health risks of indoor smoke to the weight of the cheapest stove, too heavy for most women to carry on their own. One barrier sat with suppliers: manufacturers thought distributors should run awareness campaigns while distributors thought manufacturers should, so few households heard about the health risks at all.
Our 21 recommendations fell into two areas: local access points where households could compare stoves and buy on credit, and a phone and SMS line for product information and checking that salespeople were legitimate. In the two years after the diagnostic, more than 72,000 improved cookstoves were sold in Uganda, and households using them cut their monthly fuel use by about a third (36% on average), according to World Bank reporting. Those figures cover the Bank's wider program, which our diagnostic informed.
Policy and service design around how people decide
Policy design applies what a diagnosis finds to the rules and services a government runs, and it covers the officials making the decisions as well as the public.
The Investor Office at the Ontario Securities Commission protects retail investors from fraud, partly through free education resources such as GetSmarterAboutMoney.ca. As that library grew, it became harder for visitors to know where to start. In our work with the Ontario Securities Commission, we based our redesign recommendations on research into how people learn and make financial decisions. A cluttered interface overloads readers with information, and a missing starting point can leave them stuck choosing.
We recommended organizing the existing content into courses instead of rewriting it, and sending reminders timed to real decisions. A user who said they contribute to a tax-sheltered account every year, for example, would get a message shortly before the contribution deadline. Our survey of Canadian retail investors found that the personas most likely to use the library were all saving toward a specific goal, such as retirement, so we tied the reminders to those goals.
Other policy design work happens inside government. Health Canada asked us to help refine the governance framework for organ donation and transplantation across Canada, where a steering committee with very different priorities had to agree on one model. We used interviews and surveys to define the problems. We then ran a discrete choice experiment, a survey method that measures how people trade off competing options, and followed it with a workshop where committee members reallocated points across solutions until the plan reflected the group's range of views.
Experimentation infrastructure: the capacity to learn what worked
Experimentation infrastructure is the standing set of people and routines that lets a department test a change against a control group as a normal part of rolling it out. A one-off trial answers one question, and the infrastructure lets a department keep asking. The OECD's review found the same gap on the cost side of government AI, with only about 1 in 6 UK government AI projects (16%) showing forecast costs, which leaves little to weigh benefits against.
When we help a department build this capacity, we look for four things:
- An evaluation plan written before launch, naming the outcome and the result that would change the decision.
- Data-sharing agreements signed in advance, so outcome data reaches the evaluators without a new negotiation each time.
- A default of staged or randomized rollout, so a comparison group exists whenever a change goes live.
- A shared record of results, including changes that made no difference, so the next team does not repeat them.
A randomized trial is the strongest evidence, but it is not always possible or appropriate, for example when a policy must reach everyone at once. Discrete choice experiments, like the one we ran for Health Canada, and comparisons against similar regions or earlier periods can still show which option is likely to work. We have also run controlled experiments on reducing how often older adults fall for financial scams, a question securities regulators also deal with.
The same discipline applies to evidence flowing into policy. For Nesta's Innovation Mapping Team, we reviewed methods for measuring how evidence is used in policy decisions and designed strategies to put data tools in front of policymakers at the moment they decide.
AI-assisted decision support for public servants
AI-assisted decision support helps officials weigh more evidence than a meeting can hold, and the behavioral work lies in how people inside government use what a model or scoring system tells them. Almost half of government AI cases aim at better decisions, as the OECD figures above show, so this is where behavioral science and AI meet most directly in the public sector.
Canada's National Research Council allocates hundreds of millions of dollars in federal research infrastructure funding to about 150 research facilities, each scored on 11 dimensions. Funding used to be settled in workshops where participants argued for their priorities, which rewarded bargaining over evidence. In our work with the National Research Council, we built a single composite score weighted by the council's strategic priorities.
To set the weights, we showed its director generals pairs of hypothetical facilities and asked them to pick one, collecting more than 2,500 choices over five sessions. Asking people to choose between concrete options shows priorities that a debate about principles hides. The choices showed where preferences were stable and where the group split, and a one-hour workshop of alignment exercises brought everyone to a model they accepted. We tested the system with 12 director generals, and the council's Senior Executive Committee reviewed it.
The same principle carries into larger AI programs. With BDC, Canada's bank for entrepreneurs, we worked with the Senior Management Committee and AI Council to specify ten AI capabilities in a phased roadmap, each keeping human judgment, from account managers to community partners, as the point of trust. BDC is now implementing the blueprint.
The questions we bring to any public-sector AI tool are when officials should defer to it and when they should override it, and how its interface shows uncertainty. A tool that officials trust too much spreads its errors, and one they distrust goes unused.
Frequently asked questions
What is a nudge unit?
A nudge unit is a team that applies behavioral science to public policy and tests changes with randomized trials. The UK's Behavioural Insights Team, set up in 2010, is widely described as the first central-government behavioral insights unit. Many units focus on citizen communications such as letters and reminders, and some also work on service redesign or on decisions made inside government.
What are the best alternatives to the Behavioural Insights Team and ideas42?
The best alternative depends on the work. For repeated trials on existing programs, an in-house government unit with data access often moves fastest. For problems that also need behavioral diagnosis and AI decision support, an applied behavioral science consultancy such as The Decision Lab covers more of the policy cycle within one team.
Do nudges work in government?
Yes, though the effects at scale are smaller than early academic studies suggested. A small rise across millions of people can still pay off when a change costs almost nothing to make, such as a reworded reminder letter. Problems rooted in how a service or rule works usually need the service or rule to change as well.
How can governments use AI with behavioral science?
Behavioral science shapes how officials use AI as well as what the AI recommends. That means testing when staff defer to a model or override it, and measuring results against a comparison group. Asking decision-makers to choose between concrete options, as we did with Canada's National Research Council, also turns their priorities into model weights.
What does a behavioral science firm do for a government agency?
A behavioral science firm finds out why people are not acting as a policy expects, then designs and tests changes in response. Clients range from securities regulators to development banks. The work can target citizens or the officials who run a program, such as a steering committee that needs to agree on a single governance model.

