Behavioral Science for Customer Trust in AI: Designing Transparency, Control, and Appropriate Reliance

Last updated October 9, 2026.

A Gartner survey of 5,728 customers, conducted in December 2023, found that 64% of them, about 2 in 3, would prefer companies not to use AI in customer service. 53% said they would consider switching to a competitor if they learned a company planned to use it. Their most common concern, cited by 3 in 5, was that AI would make it harder to reach a person.

At The Decision Lab, we design customer trust in AI to match two things: what the system can actually do for a given task, and what it costs the customer if the system gets that task wrong. Trust can fail in either direction. Customers who trust too little avoid help that would serve them, and customers who trust too much act on the AI's mistakes.

We have worked on this problem in health and consumer technology. With researchers at Mila, we designed the behavioral side of COVI, a contact tracing app whose AI model calculated a graded COVID-19 risk score. More recently, we built the behavioral engine behind the recommendations of a digital wellness platform serving 25 million members.

The five design levers at a glance

  1. Set expectations before the first answer. Tell customers plainly that they are dealing with AI and what it can and cannot do.
  2. Show uncertainty in a form people can act on. A graded answer with a matching next step invites better reliance than one confident answer.
  3. Give customers control. Let them adjust an answer or hand the problem to someone qualified to review it.
  4. Add friction only where mistakes are expensive. Ask customers to check before acting on high-stakes outputs, and keep low-risk tasks fast.
  5. Plan for the first visible error. People lose confidence in an algorithm faster than in a person after seeing it make a mistake, so recovery has to be designed in advance.

Appropriate reliance means customers accept AI help when it is reliable for the task, and check the answer or escalate when it is not.

What appropriate reliance means

John Lee and Katrina See's review of trust in automation, published in Human Factors in 2004, set the design goal of appropriate reliance: the trust people place in an automated system should match what the system can actually do. When trust falls below the system's capability, people ignore help that would have served them. When it rises above, they accept errors they should have caught.

For customer-facing AI, this changes what a product team should aim for. Raising trust across the board pushes customers toward overreliance in exactly the cases where the AI is wrong, so in our work we calibrate trust task by task.

The research on human and AI teams shows how hard that calibration is. Michelle Vaccaro, Abdullah Almaatouq and Thomas Malone's meta-analysis of human and AI combinations in Nature Human Behaviour pooled 106 experiments comparing people working with AI against people and AI working on their own. On average, the combinations did better than people alone but worse than whichever of the two performed best by itself. The losses were concentrated in decision tasks and in cases where the AI was more accurate than the people using it.

Matching trust to task risk

We sort the tasks a customer-facing AI handles by what happens when it is wrong, and design each tier differently.

  • Low-risk, reversible tasks, such as checking an order status: fast answers and light disclosure.
  • Moderate-risk tasks, such as explaining a policy or comparing plans: confidence cues, links to sources, edit controls and a visible route to a person.
  • High-risk or irreversible tasks, such as approving a refund exception or moving a large sum: confirmation steps, independent checks, review by a person or specialist, and an audit trail.

The same assistant can sit in all three tiers within one conversation. The design has to follow the task rather than the product.

Why customers trust AI too little

Berkeley Dietvorst, Joseph Simmons and Cade Massey's study of algorithm aversion, published in the Journal of Experimental Psychology: General in 2015, found that people who watched a forecasting algorithm make mistakes became less willing to use it than a human forecaster. This held even after they saw the algorithm outperform the human. In a customer product, the first error a customer notices can end their use of the feature, however well it performs afterwards.

Customers also judge the company behind the AI, and they judge AI involvement in ways that have little to do with output quality.

How AI disclosure affects customer trust

Oliver Schilke and Martin Reimann ran 13 experiments with more than 5,000 participants on how AI disclosure affects trust, published in Organizational Behavior and Human Decision Processes in 2025. People and organizations that disclosed their use of AI were trusted less than those that did not, in settings ranging from university teaching to investment funds. The penalty held across the different disclosure wordings they tested. Being exposed by someone else as having used AI cost more trust than disclosing it.

For product teams, concealment is therefore the riskier choice, and a label cannot carry trust on its own. We disclose plainly and early, then put our design effort into what follows the label: whether the answer proves useful, and whether a person is reachable when it does not.

Why customers trust AI too much

In November 2022, after a grandparent died, Jake Moffatt asked Air Canada's support chatbot about bereavement fares and was told the discount could be claimed after travel. The airline's actual policy required approval before the flight, and Air Canada refused the claim. Before British Columbia's Civil Resolution Tribunal, it argued it was not responsible for what its chatbot said. In Moffatt v. Air Canada, the tribunal ruled against the airline in 2024, finding it had not taken reasonable care to ensure the chatbot was accurate. It ordered Air Canada to pay the difference between the bereavement rate and the full fares.

Moffatt relied on the chatbot the way its design invited. It sat on the airline's own website and answered with confidence, with nothing to signal that it was less reliable than the policy page. The tribunal made the same point: a customer has no reason to know that one part of a company's website is accurate and another is not.

Explanations can make overreliance worse. In experiments on AI explanations and team performance presented at the 2021 ACM CHI Conference on Human Factors in Computing Systems, Gagan Bansal and colleagues found that explanations raised the chance that people accepted the AI's recommendation whether or not it was correct. The explanations also did no better than simply showing how confident the AI was. Generative AI produces fluent explanations by default, so customer products built on it carry this risk from the start.

1. Set expectations before the first answer

Say that the customer is dealing with AI, and what it is for, at the start of the interaction rather than in a footnote. Given Schilke and Reimann's finding that exposure costs more trust than disclosure, concealment is a poor bet. Scope matters as much as the label. A support assistant that says it can explain policies but cannot approve exceptions gives customers a reason to check before relying on a promise.

In our published applied research with Mila on COVI, comprehension sat at the center of the design. We ran weekly surveys of Canadians during the pandemic, measuring trust in tracing apps and privacy concerns among other attitudes. We built the results into a five-pillar decision model covering user preference, comprehension, wellbeing, empowerment and inclusivity. Members of our team also co-authored a paper with Mila researchers in The Lancet Digital Health on the need for privacy in public digital contact tracing.

2. Show uncertainty in a form people can act on

A single confident answer invites the same reliance whether the system is right or wrong. For customers, uncertainty works best as a graded output, where what the system is sure about and what it is unsure about each come with a next step. Bansal and colleagues' finding that showing confidence matched elaborate explanations supports starting there.

In practice, a support assistant that would otherwise say "You are eligible for a refund" can say: "Based on the details provided, you may be eligible. Before booking, review the policy conditions below or ask a support specialist to confirm." The customer gets a useful answer, a reason to check, and a route to a person, and the answer does not feel evasive.

At the time of the COVI project, contact tracing apps could tell people only whether they had been exposed or not. We designed graded risk guidance in place of those blunt alerts, with messaging matched to the user's level of risk.

3. Give customers control over the AI's output

In a 2018 follow-up in Management Science, Dietvorst, Simmons and Massey found that people were more likely to use an imperfect algorithm when they could adjust its forecasts, even slightly. In customer products, control takes two forms: adjusting what the AI produces, and leaving the AI for meaningful review. Both need to be visible before the customer needs them. An escalation route that appears only after several failed attempts feels like a trap to the customer who finally finds it.

4. Add friction only where mistakes are expensive

Zana Buçinca, Maja Barbara Malaya and Krzysztof Gajos tested cognitive forcing designs, such as asking people to make their own decision before seeing the AI's recommendation, in a 2021 experiment with 199 participants. Compared with simple explanations, these designs reduced overreliance. Participants also rated the designs that reduced overreliance most as the least favorable, and the benefits went mainly to people who enjoy effortful thinking.

We place friction at the high-risk tier, where acting on a wrong answer is costly or hard to reverse. On low-risk tasks, friction mostly teaches customers to avoid the product. Because the benefits in the Buçinca study were uneven, we test these safeguards with customers who differ in expertise and digital confidence. A check that protects an expert customer can shut out someone less confident with digital tools.

5. Plan for the first visible error

Because of algorithm aversion, the first error a customer notices does disproportionate damage. We work out in advance how the product acknowledges a mistake and what it offers to put it right, including a direct route to review. The Air Canada ruling adds a practical requirement. The AI's answers have to agree with the company's own published policies, because the company is responsible for both.

How we measure customer trust in AI

Stated trust and reliance are different measures, and we track both. Surveys tell us how customers feel about the AI. To measure reliance, we look at behavior separately for correct and incorrect AI outputs. Appropriate reliance improves when customers accept correct outputs more often than incorrect ones, without unnecessary abandonment or escalation on low-risk tasks where the AI is confident.

In practice, we track:

  • the acceptance rate when the AI is correct
  • the acceptance rate when the AI is incorrect
  • override and escalation rates
  • task completion
  • abandonment after a visible error
  • each of these broken down by customer segment and task risk tier.

Segmenting customers matters because attitudes to a product vary more than demographics suggest. In a client engagement with a leading online learning platform, published as an anonymized case study, we surveyed more than 800 students about how they use study tools as AI changes how they learn. Willingness to subscribe, an attitude, produced a cleaner and more predictive segmentation than demographics did. We segment customers by how they approach an AI product in the same way before designing for them.

We also test trust designs before launch. In a client engagement with a global consumer electronics company, every target behavior for a new mindfulness service, built into a health app used by 60 million people, was tested through experimentation before any production code was written. Design time fell by three quarters. In our engagement with the digital wellness platform, we ran the full recommendation engine against synthetic members in a validation harness before it reached the product. That let the team find and fix design failures during design rather than after launch.

Frequently asked questions

What is appropriate reliance on AI?

Appropriate reliance means people trust an AI system as much as its performance on a given task justifies. They accept its output where it is reliable and check or escalate where it is not. The term comes from Lee and See's 2004 review of trust in automation. For customer products, we also weigh what being wrong would cost the customer.

Does labeling advice as AI-generated reduce customer trust?

Usually, yes. Across 13 experiments, Schilke and Reimann found that disclosing AI use lowered trust under every wording they tested, and being exposed by someone else lowered it further. We recommend disclosing plainly at the start, then designing what comes after the label, including visible confidence levels and an easy route to a person.

Do AI explanations help customers trust AI appropriately?

On their own, they often do not. In experiments by Bansal and colleagues, explanations made people more likely to accept an AI's recommendation whether it was right or wrong, and they did no better than showing the AI's confidence. Explanations help most when paired with clear uncertainty signals and a concrete reason to check.

How do you measure customer trust in an AI product?

We pair surveys of stated trust with behavioral measures, tracking acceptance separately for correct and incorrect AI outputs. Reliance is improving when customers accept correct outputs more often than incorrect ones, without dropping off on low-risk tasks. We also track escalation, task completion and abandonment after errors, broken down by customer segment and task risk.

Should customers always be able to reach a human when using AI?

For consequential decisions, customers need a clearly visible path to meaningful review before they rely on a high-stakes AI output. In many customer service settings, that means reaching a person. In others, it may mean a qualified specialist, a formal appeals process, an emergency channel or another way to complete the task.

When should a product add friction to an AI interaction?

Add friction when an AI-supported action is costly or hard to reverse, such as approving a refund exception, transferring a large sum, changing medical information or making a legally binding commitment. On low-risk tasks, friction tends to reduce use without improving decisions, so we match each safeguard to what the customer stands to lose.

Sources

  1. Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., & Weld, D. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. CHI '21.
  2. Bengio, Y., Janda, R., Yu, Y. W., Ippolito, D., Jarvie, M., Pilat, D., Struck, B., Krastev, S., & Sharma, A. (2020). The need for privacy with public digital contact tracing during the COVID-19 pandemic. The Lancet Digital Health.
  3. Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction.
  4. Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114-126.
  5. Dietvorst, B. J., Simmons, J. P., & Massey, C. (2018). Overcoming algorithm aversion: People will use imperfect algorithms if they can (even slightly) modify them. Management Science, 64(3), 1155-1170.
  6. Gartner (2024). Survey of 5,728 customers conducted December 2023 on AI in customer service.
  7. Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1).
  8. Moffatt v. Air Canada, British Columbia Civil Resolution Tribunal (2024).
  9. Schilke, O., & Reimann, M. (2025). The transparency dilemma: How AI disclosure erodes trust. Organizational Behavior and Human Decision Processes.
  10. Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour.
  11. The Decision Lab. Changing how billions of phones communicate risk (COVI case study, published applied research).
  12. The Decision Lab. Building the behavioral engine inside a digital wellness platform (anonymized client case study).
  13. The Decision Lab. Reducing product feature design time using simulation-first validation (anonymized client case study).
  14. The Decision Lab. Keeping students learning in the age of AI through behavioral segmentation (anonymized client case study).
  15. The Decision Lab. Algorithm aversion (reference guide).

About the Authors

A man in a blue, striped shirt smiles while standing indoors, surrounded by green plants and modern office decor.

Dan Pilat

Managing Director

Dan is a Co-Founder and Managing Director at The Decision Lab. He is a bestselling author of Intention - a book he wrote with Wiley on the mindful application of behavioral science in organizations. Dan has a background in organizational decision making, with a BComm in Decision & Information Systems from McGill University. He has worked on enterprise-level behavioral architecture at TD Securities and BMO Capital Markets, where he advised management on the implementation of systems processing billions of dollars per week. Driven by an appetite for the latest in technology, Dan created a course on business intelligence and lectured at McGill University, and has applied behavioral science to topics such as augmented and virtual reality.

A smiling man stands in an office, wearing a dark blazer and black shirt, with plants and glass-walled rooms in the background.

Dr. Sekoul Krastev

Managing Director & Co-Founder

Dr. Sekoul Krastev is a decision scientist and Co-Founder of The Decision Lab, one of the world's leading behavioral science consultancies. His team works with large organizations—Fortune 500 companies, governments, foundations and supernationals—to apply behavioral science and decision theory for social good. He holds a PhD in neuroscience from McGill University and is currently a visiting scholar at NYU. His work has been featured in academic journals as well as in The New York Times, Forbes, and Bloomberg. He is also the author of Intention (Wiley, 2024), a bestselling book on the science of human agency. Before founding The Decision Lab, he worked at the Boston Consulting Group and Google.

Notes illustration

Eager to learn about how behavioral science can help your organization?