Why Users Stop Engaging With Health and Wellness Apps: A Behavioral Checklist

Last updated October 9, 2026

Five behavioral failure modes in rewards, goals, recommendations and prompt timing, with tests to identify what is driving disengagement.

Health and wellness apps lose users when their rewards and recommendations are built around what the platform can count, and not around how each person decides what to do next. At The Decision Lab, we design the incentive and personalization logic inside digital health platforms. Most recently, in a confidential engagement with a digital wellness platform serving 25 million members, we built the behavioral engine that now powers its production recommender.

Across that work, we see engagement stall for five behavioral reasons: rewards track the wrong behavior, one reward structure is applied to everyone, external rewards fade before habits form, personalization runs without behavioral logic, and prompts arrive at the wrong moment. This checklist pairs each one with a test your team can run on its own product and the design change we would make. It applies to consumer health apps and employer wellness platforms alike, and to any product that uses points or personalized nudges to change what people do.

In this guide, a behavioral engine is the decision logic inside a health platform that determines which action, goal, reward or prompt is most likely to help a specific user make progress. Each failure mode below is a gap in that logic.

The five-point checklist

  1. Rewards track the wrong behavior. Check whether the actions you reward are linked to the outcome you want, or only to activity your platform can log.
  2. One reward structure for everyone. Check whether reward values and goals adapt to each user's readiness, or whether already-active users collect most of the points.
  3. External rewards fade before habits form. Check whether rewards hand off gradually to routines users keep on their own once the reward ends.
  4. Personalization without behavioral logic. Check whether you can explain why each user receives each recommendation, beyond matching it to their profile data.
  5. Prompts arrive at the wrong moment. Check whether prompts reach users when they are able to act, or on a fixed schedule.

1. Rewards track the wrong behavior

Most wellness platforms reward what their systems can see, such as daily logins and completed challenges. Those actions are easy to count, but a user can rack them up without changing anything that affects their health.

We keep four kinds of outcomes apart in every engagement project, because a platform can improve one without moving the next:

Outcome type What it measures Examples
Engagement Use of the product Opens, return rate, completion, retention
Behavior What users do in their lives Activity, adherence, attendance, follow-through
Clinical Health status Blood pressure, symptom scores, lab markers, body weight
Commercial Value to the buyer Program use, sponsor retention, cost avoidance, renewals

The gap between those layers shows up in the strongest evidence on employer wellness programs. In a randomized trial covering 32,974 employees at BJ's Wholesale Club, Zirui Song and Katherine Baicker found that worksites offering a wellness program had a share of employees reporting regular exercise about 8 percentage points higher than control worksites, a self-reported behavior outcome. After 18 months, the program made no significant difference to 10 clinical markers, including blood pressure and cholesterol, or to health care spending.

We treat a reward system as a statement of what the platform values. In the behavioral engine we built for the 25-million-member wellness platform, the points an action earns equal its clinical weight times its difficulty, adjusted for the sponsor's priorities. An action that matters more for health and asks more of the member is worth more.

Test: export the 10 actions that earned the most points last quarter. For each, record points paid, users reached, completion rate and the evidence linking the action to the behavior or clinical outcome your sponsor values. If most points went to actions with no such evidence, the reward system is paying for engagement alone.

What we change: we weight rewards by how strongly the evidence links each action to health, adjusted for the effort it asks of the user. We reward progress toward a behavior or clinical outcome alongside task completion, and keep a few easy actions as entry points for new users.

2. One reward structure for everyone

A single set of point values and goals treats a marathon runner and someone who has not exercised in a year as the same user. The runner collects rewards for what they would have done anyway, and the person who most needs support faces goals that feel out of reach.

The same incentive does very different work depending on who receives it. Gary Charness and Uri Gneezy paid university students to attend the gym eight times in a month and found clear increases in attendance after the payments ended, compared with control groups. The whole increase came from people who had not been regular gym-goers before the study, so for the regulars the payments changed nothing that lasted.

The 25-million-member wellness platform we worked with was running its recommendations on a single set of assumptions applied to every member, regardless of who they were or where they were in their journey. We replaced it with a behavioral engine that keeps a behavioral profile for each member and updates it as their behavior changes, so the same action can carry a different reward and a different place in the queue for two members.

Test: group users by how active they were in their first two weeks. For each group, pull its share of users, its share of points earned, its 90-day retention and its change on your main behavior or clinical measure. If the most active fifth earns most of the points and shows the least change, the reward structure is paying the wrong people.

What we change: we group users by their past behavior and readiness to change, then set reward values and goal difficulty for each group. Demographic segments are a starting point at most.

3. External rewards fade before habits form

Points and badges can lift use quickly, and that lift often disappears once the reward stops or loses its novelty. The habit the program was meant to build never had time to take over from the reward.

Heather Royer, Mark Stehr and Justin Sydnor tested this with about 1,000 employees at a Fortune 500 company, paying $10 per visit to the company gym over four weeks. The payments left only small lasting effects on gym use. Employees who were also offered a commitment contract at the end of the program, putting up their own money against their exercise goal, showed significant changes that were still detectable several years after the payments ended.

Rewards also fail when the goals behind them set users up to fall short. For the MyHeart Counts Canada fitness app, we ran the user research and designed the behavioral interventions. The research showed that guilt over missed goals pushed people to avoid the app and eventually uninstall it. We recommended asking users to report their current activity level right before they set a goal, so their targets stay close to existing habits while still asking a little more of them.

Test: for 12 weeks after a reward ends, pull two weekly measures for users who earned it: active use of the app and completion of the rewarded behavior. Pull the same measures for similar users who never received the reward. If the two groups converge on both measures within a month, the reward created a spike and no habit.

What we change: we design rewards to step down gradually and hand over to the user's own commitments, and we set early goals that most users can meet.

4. Personalization without behavioral logic

Most health platforms already collect enough data to personalize. What they usually lack is a behavioral engine, the rules for what to change for each person and why. Without one, the data shapes greetings and content feeds while the recommendations stay the same for everyone.

A member who keeps saving workout videos but never starts one probably does not need another exercise recommendation. They may need a smaller first step, or a prompt at a time of day when they have room to act.

A recommendation can be clinically correct and still go nowhere, because whether someone acts on it depends on barriers the clinical logic never sees. The behavioral engine we built for the 25-million-member wellness platform scores every candidate action for each member against personal barriers, starting with how much effort it takes them to get started. It uses that score to decide what surfaces for a member today, and it adjusts as the member changes.

Personalization can fail in ways that only show up at scale, so we tested the engine before it reached a single real member. We ran it against simulated members in a test environment, and design flaws surfaced in an afternoon instead of six months into production.

Test: choose your three most common recommendations. For each, pull how many users received it, opened it, accepted it and completed the underlying behavior, broken down by user group. If the same recommendation reaches users with very different barriers and few of them complete it in any group, your system is matching content to profile data. A behavioral engine matches each action to the barriers of the person receiving it.

What we change: we write down the rules behind each recommendation as part of the behavioral engine and test them on simulated users before launch. The client owns the engine's specification, so their team can add actions and member groups without a rebuild.

5. Prompts arrive at the wrong moment

Even a well-matched reward or recommendation fails if it arrives when a user cannot act on it. A reminder sent at 9 a.m. every day reaches some users in a meeting and others on a walk they were already taking.

The HeartSteps study by Predrag Klasnja and colleagues shows both the value of timing and its limits. In the full six-week analysis, walking and activity suggestions tailored to each person's context were followed by 14% more steps over the next 30 minutes than no suggestion, about 35 extra steps on a typical 253, though this average effect fell just short of conventional statistical significance. Suggestions went out at times the participants chose, with each one randomized. The effect was far larger at the start of the study, 66% more steps, and shrank as people got used to the prompts.

We apply the same principle outside health. In our work with the Ontario Securities Commission, we developed evidence-based design recommendations for its investor education website, including a survey of Canadian retail investors. The research was clear that investor information has the most effect when it reaches people just before a key decision, and personalization is what makes that timing possible. In the behavioral engine we built for the wellness platform, which prompt fires and when is set by the same logic that sets rewards.

Test: for each notification type, pull the send time, the weeks since the user signed up, whether the user acted within an hour and whether they turned notifications off afterward. A flat response across send times means you are sending on a schedule, and a falling response across weeks means users are tuning the prompts out.

What we change: we tie prompts to moments when the user can act, such as a gap in their day or a point in the app where a decision is made, and we reduce how often a prompt repeats as its effect wears off.

How we rebuild engagement in a health platform

We start every engagement project by finding which of the five failure modes is holding back which users, because the fix for a timing problem does little for a reward that pays the wrong behavior. Our projects usually run in four stages:

  1. Diagnostic immersion. We audit how the platform's current rewards and prompts affect behavior, and map the highest-impact opportunities for redesign.
  2. Evidence synthesis. We review research on behavior change and clinical evidence, along with what has worked and failed on platforms in nearby fields such as gaming and personal finance, and map it onto established behavior change frameworks.
  3. Member and employer research. We pair member surveys with AI-powered focus groups to learn how members experience the current mechanics and what would make them meaningful. Our guide to behavioral research methods, from surveys to simulations covers how we choose between them.
  4. Framework design and specification. We build an evidence-based behavior change framework and turn it into specifications the client's engineers can build from.

Our behavioral engine for a digital wellness platform shows the end result. In that confidential engagement, the platform served 25 million members and was running every recommendation on a single set of assumptions. We built the behavioral engine that prices each action in points and decides which recommendation and prompt a member gets each day, adjusting as the member changes. It now powers the platform's production recommender, and because the client owns its specification, their team extends it to new actions and member groups on its own.

If your platform has enough member data to personalize but cannot explain why one user gets a given action or reward and another user does not, the missing layer is usually a behavioral engine. We help teams diagnose that logic and test it before launch, then turn it into specifications their product and engineering teams can maintain.

When to bring in a behavioral design partner

A product or UX partner is often the right starting point when users cannot find features or complete essential tasks. When the interface works but engagement has plateaued, the problem usually sits deeper, in the reward, goal, recommendation and timing logic that shapes repeated behavior. In those cases, rebuilding engagement means redesigning the behavioral logic behind the experience as well as the interface.

The signs we see most often before a team calls us:

  • New features ship, and weekly active use stays flat.
  • Employer or health plan buyers ask for clinical outcomes, and the platform can only report activity.
  • Reward costs keep rising while the same small group of users earns most of the points.
  • The data and infrastructure for personalization exist, but recommendations look the same for most users.

Our team combines behavioral science research with product design, so we can diagnose why engagement stalls and also specify the behavioral engine that replaces the old logic. We describe how we approach these projects on our page on behavioral design for health and wellness platforms, and our guide to behavioral science for health app design covers the wider design principles. For buyers weighing the evidence base itself, we also answer whether behavioral science consulting is evidence-based.

Frequently asked questions

Why do users stop using health apps?

Independent usage data shows most people stop opening popular mental health apps within a month of installing them. In our work, the most common causes are rewards that pay for easy activity and goals that do not fit the person, and poorly timed prompts make both worse. Each cause needs a different design change, so the first step is finding which one affects which group of users.

Does gamification work in health apps?

Gamification lifts engagement quickly, but points and badges on their own rarely build lasting habits. Field experiments that paid people to exercise found strong effects while the payments ran and small effects afterward, unless the program also helped people commit on their own. Game mechanics work best when they reward outcome-linked behaviors and step down over time.

How do you personalize incentives in a wellness program?

Start by modeling each user's past behavior and barriers, such as how much effort an action takes them to begin. Set reward values and goal difficulty for each group of users, and price actions by how strongly the evidence links them to health. Test the rules on simulated users before launch, so design flaws surface before real members meet them.

How do you measure whether app engagement improves health outcomes?

Measure engagement and health separately. Engagement covers opens and retention, while health covers behaviors such as activity and clinical markers such as blood pressure. Compare users who received a new feature with a similar group who did not. Large employer wellness trials show engagement can rise with no change in health, and a randomized rollout gives the clearest answer when feasible.

What is a behavioral engine?

A behavioral engine is the decision logic inside a health platform that determines which action, goal, reward or prompt is most likely to help a specific user make progress. It models each user's behavior and barriers and updates as they change. We built one for a wellness platform serving 25 million members, and it now powers that platform's production recommender.

Sources

  1. Baumel, A., Muench, F., Edan, S., & Kane, J. M. (2019). Objective user engagement with mental health apps: Systematic search and panel-based usage analysis. Journal of Medical Internet Research, 21(9), e14567.
  2. Charness, G., & Gneezy, U. (2009). Incentives to exercise. Econometrica, 77(3), 909-931.
  3. Klasnja, P., et al. (2019). Efficacy of contextually tailored suggestions for physical activity: A micro-randomized optimization trial of HeartSteps. Annals of Behavioral Medicine.
  4. Royer, H., Stehr, M., & Sydnor, J. (2015). Incentives, commitments, and habit formation in exercise: Evidence from a field experiment with workers at a Fortune-500 company. American Economic Journal: Applied Economics, 7(3), 51-84.
  5. Song, Z., & Baicker, K. (2019). Effect of a workplace wellness program on employee health and economic outcomes: A randomized clinical trial. JAMA, 321(15), 1491-1501.
  6. The Decision Lab. Building the behavioral engine inside a digital wellness platform serving 25 million members (case study).
  7. The Decision Lab. Designing a cutting-edge health app to get Canadians moving (case study).
  8. The Decision Lab. Protecting Canadian investors from fraud (case study).

About the Authors

A man in a blue, striped shirt smiles while standing indoors, surrounded by green plants and modern office decor.

Dan Pilat

Managing Director

Dan is a Co-Founder and Managing Director at The Decision Lab. He is a bestselling author of Intention - a book he wrote with Wiley on the mindful application of behavioral science in organizations. Dan has a background in organizational decision making, with a BComm in Decision & Information Systems from McGill University. He has worked on enterprise-level behavioral architecture at TD Securities and BMO Capital Markets, where he advised management on the implementation of systems processing billions of dollars per week. Driven by an appetite for the latest in technology, Dan created a course on business intelligence and lectured at McGill University, and has applied behavioral science to topics such as augmented and virtual reality.

A smiling man stands in an office, wearing a dark blazer and black shirt, with plants and glass-walled rooms in the background.

Dr. Sekoul Krastev

Managing Director & Co-Founder

Dr. Sekoul Krastev is a decision scientist and Co-Founder of The Decision Lab, one of the world's leading behavioral science consultancies. His team works with large organizations—Fortune 500 companies, governments, foundations and supernationals—to apply behavioral science and decision theory for social good. He holds a PhD in neuroscience from McGill University and is currently a visiting scholar at NYU. His work has been featured in academic journals as well as in The New York Times, Forbes, and Bloomberg. He is also the author of Intention (Wiley, 2024), a bestselling book on the science of human agency. Before founding The Decision Lab, he worked at the Boston Consulting Group and Google.

Notes illustration

Eager to learn about how behavioral science can help your organization?