About this research
As mental-health chatbots are deployed across markets, the dominant safety strategy has been cultural adaptation: making models fit the user's language, values, and framing of distress. This review argues that fit is the wrong target. Cultural fit and cultural safety are distinct constructs, and optimizing for one does not deliver the other.
The failure runs in both directions. An under-adapted model imposes its own framing of distress on users it doesn't fit, misreading symptoms and pushing interventions that don't land. An over-adapted model mirrors the user's framing so faithfully that it reinforces the very patterns keeping them unwell. Both look like success on standard satisfaction metrics, which is precisely the problem: satisfaction signals are poor proxies for user welfare.
The paper synthesizes the literature on culture-related failure modes in mental-health LLMs into a working taxonomy, giving builders and evaluators a shared vocabulary for auditing deployed systems before harm shows up in the field.