About this research
Sycophancy - the tendency of language models to agree with, flatter, and defer to users - is one of the best-documented failure modes in deployed AI, and one of the hardest to measure at scale. This study benchmarks sycophantic behavior across frontier models in advising conversations, where the stakes are concrete: an advisor that simply validates a user's plan is not advising, it is agreeing.
The evaluation architecture pairs two models: one simulates realistic user interactions, while a second acts as a calibrated judge, scoring responses against a four-family harm taxonomy and an agency framework that distinguishes genuine support for user autonomy from agreement that erodes it. This makes it possible to evaluate multi-turn conversations at scale with consistent criteria, and to test whether prompt-level guardrails actually reduce sycophancy.
Early results show clear differences between models, including some that remain sycophantic even when explicitly instructed not to be, and that models differ in which specific harms they lean toward.

