Last updated October 9, 2026
Five behavioral failure modes in rewards, goals, recommendations and prompt timing, with tests to identify what is driving disengagement.
Health and wellness apps lose users when their rewards and recommendations are built around what the platform can count, and not around how each person decides what to do next. At The Decision Lab, we design the incentive and personalization logic inside digital health platforms. Most recently, in a confidential engagement with a digital wellness platform serving 25 million members, we built the behavioral engine that now powers its production recommender.
Across that work, we see engagement stall for five behavioral reasons: rewards track the wrong behavior, one reward structure is applied to everyone, external rewards fade before habits form, personalization runs without behavioral logic, and prompts arrive at the wrong moment. This checklist pairs each one with a test your team can run on its own product and the design change we would make. It applies to consumer health apps and employer wellness platforms alike, and to any product that uses points or personalized nudges to change what people do.
In this guide, a behavioral engine is the decision logic inside a health platform that determines which action, goal, reward or prompt is most likely to help a specific user make progress. Each failure mode below is a gap in that logic.
The five-point checklist
- Rewards track the wrong behavior. Check whether the actions you reward are linked to the outcome you want, or only to activity your platform can log.
- One reward structure for everyone. Check whether reward values and goals adapt to each user's readiness, or whether already-active users collect most of the points.
- External rewards fade before habits form. Check whether rewards hand off gradually to routines users keep on their own once the reward ends.
- Personalization without behavioral logic. Check whether you can explain why each user receives each recommendation, beyond matching it to their profile data.
- Prompts arrive at the wrong moment. Check whether prompts reach users when they are able to act, or on a fixed schedule.
1. Rewards track the wrong behavior
Most wellness platforms reward what their systems can see, such as daily logins and completed challenges. Those actions are easy to count, but a user can rack them up without changing anything that affects their health.
We keep four kinds of outcomes apart in every engagement project, because a platform can improve one without moving the next:
| Outcome type | What it measures | Examples |
|---|---|---|
| Engagement | Use of the product | Opens, return rate, completion, retention |
| Behavior | What users do in their lives | Activity, adherence, attendance, follow-through |
| Clinical | Health status | Blood pressure, symptom scores, lab markers, body weight |
| Commercial | Value to the buyer | Program use, sponsor retention, cost avoidance, renewals |
The gap between those layers shows up in the strongest evidence on employer wellness programs. In a randomized trial covering 32,974 employees at BJ's Wholesale Club, Zirui Song and Katherine Baicker found that worksites offering a wellness program had a share of employees reporting regular exercise about 8 percentage points higher than control worksites, a self-reported behavior outcome. After 18 months, the program made no significant difference to 10 clinical markers, including blood pressure and cholesterol, or to health care spending.
We treat a reward system as a statement of what the platform values. In the behavioral engine we built for the 25-million-member wellness platform, the points an action earns equal its clinical weight times its difficulty, adjusted for the sponsor's priorities. An action that matters more for health and asks more of the member is worth more.
Test: export the 10 actions that earned the most points last quarter. For each, record points paid, users reached, completion rate and the evidence linking the action to the behavior or clinical outcome your sponsor values. If most points went to actions with no such evidence, the reward system is paying for engagement alone.
What we change: we weight rewards by how strongly the evidence links each action to health, adjusted for the effort it asks of the user. We reward progress toward a behavior or clinical outcome alongside task completion, and keep a few easy actions as entry points for new users.
2. One reward structure for everyone
A single set of point values and goals treats a marathon runner and someone who has not exercised in a year as the same user. The runner collects rewards for what they would have done anyway, and the person who most needs support faces goals that feel out of reach.
The same incentive does very different work depending on who receives it. Gary Charness and Uri Gneezy paid university students to attend the gym eight times in a month and found clear increases in attendance after the payments ended, compared with control groups. The whole increase came from people who had not been regular gym-goers before the study, so for the regulars the payments changed nothing that lasted.
The 25-million-member wellness platform we worked with was running its recommendations on a single set of assumptions applied to every member, regardless of who they were or where they were in their journey. We replaced it with a behavioral engine that keeps a behavioral profile for each member and updates it as their behavior changes, so the same action can carry a different reward and a different place in the queue for two members.
Test: group users by how active they were in their first two weeks. For each group, pull its share of users, its share of points earned, its 90-day retention and its change on your main behavior or clinical measure. If the most active fifth earns most of the points and shows the least change, the reward structure is paying the wrong people.
What we change: we group users by their past behavior and readiness to change, then set reward values and goal difficulty for each group. Demographic segments are a starting point at most.
3. External rewards fade before habits form
Points and badges can lift use quickly, and that lift often disappears once the reward stops or loses its novelty. The habit the program was meant to build never had time to take over from the reward.
Heather Royer, Mark Stehr and Justin Sydnor tested this with about 1,000 employees at a Fortune 500 company, paying $10 per visit to the company gym over four weeks. The payments left only small lasting effects on gym use. Employees who were also offered a commitment contract at the end of the program, putting up their own money against their exercise goal, showed significant changes that were still detectable several years after the payments ended.
Rewards also fail when the goals behind them set users up to fall short. For the MyHeart Counts Canada fitness app, we ran the user research and designed the behavioral interventions. The research showed that guilt over missed goals pushed people to avoid the app and eventually uninstall it. We recommended asking users to report their current activity level right before they set a goal, so their targets stay close to existing habits while still asking a little more of them.
Test: for 12 weeks after a reward ends, pull two weekly measures for users who earned it: active use of the app and completion of the rewarded behavior. Pull the same measures for similar users who never received the reward. If the two groups converge on both measures within a month, the reward created a spike and no habit.
What we change: we design rewards to step down gradually and hand over to the user's own commitments, and we set early goals that most users can meet.
4. Personalization without behavioral logic
Most health platforms already collect enough data to personalize. What they usually lack is a behavioral engine, the rules for what to change for each person and why. Without one, the data shapes greetings and content feeds while the recommendations stay the same for everyone.
A member who keeps saving workout videos but never starts one probably does not need another exercise recommendation. They may need a smaller first step, or a prompt at a time of day when they have room to act.
A recommendation can be clinically correct and still go nowhere, because whether someone acts on it depends on barriers the clinical logic never sees. The behavioral engine we built for the 25-million-member wellness platform scores every candidate action for each member against personal barriers, starting with how much effort it takes them to get started. It uses that score to decide what surfaces for a member today, and it adjusts as the member changes.
Personalization can fail in ways that only show up at scale, so we tested the engine before it reached a single real member. We ran it against simulated members in a test environment, and design flaws surfaced in an afternoon instead of six months into production.
Test: choose your three most common recommendations. For each, pull how many users received it, opened it, accepted it and completed the underlying behavior, broken down by user group. If the same recommendation reaches users with very different barriers and few of them complete it in any group, your system is matching content to profile data. A behavioral engine matches each action to the barriers of the person receiving it.
What we change: we write down the rules behind each recommendation as part of the behavioral engine and test them on simulated users before launch. The client owns the engine's specification, so their team can add actions and member groups without a rebuild.
5. Prompts arrive at the wrong moment
Even a well-matched reward or recommendation fails if it arrives when a user cannot act on it. A reminder sent at 9 a.m. every day reaches some users in a meeting and others on a walk they were already taking.
The HeartSteps study by Predrag Klasnja and colleagues shows both the value of timing and its limits. In the full six-week analysis, walking and activity suggestions tailored to each person's context were followed by 14% more steps over the next 30 minutes than no suggestion, about 35 extra steps on a typical 253, though this average effect fell just short of conventional statistical significance. Suggestions went out at times the participants chose, with each one randomized. The effect was far larger at the start of the study, 66% more steps, and shrank as people got used to the prompts.
We apply the same principle outside health. In our work with the Ontario Securities Commission, we developed evidence-based design recommendations for its investor education website, including a survey of Canadian retail investors. The research was clear that investor information has the most effect when it reaches people just before a key decision, and personalization is what makes that timing possible. In the behavioral engine we built for the wellness platform, which prompt fires and when is set by the same logic that sets rewards.
Test: for each notification type, pull the send time, the weeks since the user signed up, whether the user acted within an hour and whether they turned notifications off afterward. A flat response across send times means you are sending on a schedule, and a falling response across weeks means users are tuning the prompts out.
What we change: we tie prompts to moments when the user can act, such as a gap in their day or a point in the app where a decision is made, and we reduce how often a prompt repeats as its effect wears off.
How we rebuild engagement in a health platform
We start every engagement project by finding which of the five failure modes is holding back which users, because the fix for a timing problem does little for a reward that pays the wrong behavior. Our projects usually run in four stages:
- Diagnostic immersion. We audit how the platform's current rewards and prompts affect behavior, and map the highest-impact opportunities for redesign.
- Evidence synthesis. We review research on behavior change and clinical evidence, along with what has worked and failed on platforms in nearby fields such as gaming and personal finance, and map it onto established behavior change frameworks.
- Member and employer research. We pair member surveys with AI-powered focus groups to learn how members experience the current mechanics and what would make them meaningful. Our guide to behavioral research methods, from surveys to simulations covers how we choose between them.
- Framework design and specification. We build an evidence-based behavior change framework and turn it into specifications the client's engineers can build from.
Our behavioral engine for a digital wellness platform shows the end result. In that confidential engagement, the platform served 25 million members and was running every recommendation on a single set of assumptions. We built the behavioral engine that prices each action in points and decides which recommendation and prompt a member gets each day, adjusting as the member changes. It now powers the platform's production recommender, and because the client owns its specification, their team extends it to new actions and member groups on its own.
If your platform has enough member data to personalize but cannot explain why one user gets a given action or reward and another user does not, the missing layer is usually a behavioral engine. We help teams diagnose that logic and test it before launch, then turn it into specifications their product and engineering teams can maintain.
When to bring in a behavioral design partner
A product or UX partner is often the right starting point when users cannot find features or complete essential tasks. When the interface works but engagement has plateaued, the problem usually sits deeper, in the reward, goal, recommendation and timing logic that shapes repeated behavior. In those cases, rebuilding engagement means redesigning the behavioral logic behind the experience as well as the interface.
The signs we see most often before a team calls us:
- New features ship, and weekly active use stays flat.
- Employer or health plan buyers ask for clinical outcomes, and the platform can only report activity.
- Reward costs keep rising while the same small group of users earns most of the points.
- The data and infrastructure for personalization exist, but recommendations look the same for most users.
Our team combines behavioral science research with product design, so we can diagnose why engagement stalls and also specify the behavioral engine that replaces the old logic. We describe how we approach these projects on our page on behavioral design for health and wellness platforms, and our guide to behavioral science for health app design covers the wider design principles. For buyers weighing the evidence base itself, we also answer whether behavioral science consulting is evidence-based.
Frequently asked questions
Why do users stop using health apps?
Independent usage data shows most people stop opening popular mental health apps within a month of installing them. In our work, the most common causes are rewards that pay for easy activity and goals that do not fit the person, and poorly timed prompts make both worse. Each cause needs a different design change, so the first step is finding which one affects which group of users.
Does gamification work in health apps?
Gamification lifts engagement quickly, but points and badges on their own rarely build lasting habits. Field experiments that paid people to exercise found strong effects while the payments ran and small effects afterward, unless the program also helped people commit on their own. Game mechanics work best when they reward outcome-linked behaviors and step down over time.
How do you personalize incentives in a wellness program?
Start by modeling each user's past behavior and barriers, such as how much effort an action takes them to begin. Set reward values and goal difficulty for each group of users, and price actions by how strongly the evidence links them to health. Test the rules on simulated users before launch, so design flaws surface before real members meet them.
How do you measure whether app engagement improves health outcomes?
Measure engagement and health separately. Engagement covers opens and retention, while health covers behaviors such as activity and clinical markers such as blood pressure. Compare users who received a new feature with a similar group who did not. Large employer wellness trials show engagement can rise with no change in health, and a randomized rollout gives the clearest answer when feasible.
What is a behavioral engine?
A behavioral engine is the decision logic inside a health platform that determines which action, goal, reward or prompt is most likely to help a specific user make progress. It models each user's behavior and barriers and updates as they change. We built one for a wellness platform serving 25 million members, and it now powers that platform's production recommender.

