Measuring Sycophancy in LLM Advisors: An LLM-as-Judge Framework with Simulated Users

A benchmark of sycophantic behavior across frontier models in advising contexts, built on a two-model evaluation architecture: one model simulates realistic user interactions while a second acts as a calibrated judge. Responses are scored against a four-family harm taxonomy and an agency framework separating genuine support from harmful agreement. The study also tests whether prompt-level guardrails reduce sycophancy, finding that some models remain sycophantic even when explicitly instructed otherwise.

Current focus

The full conversation dataset has been generated; analysis is underway on model-level differences, harm-type predispositions, and the effectiveness of prompt-level guardrails.

Collaborators

  • The Decision LabApplied research team
  • MilaResearch partner

About this research

Sycophancy - the tendency of language models to agree with, flatter, and defer to users - is one of the best-documented failure modes in deployed AI, and one of the hardest to measure at scale. This study benchmarks sycophantic behavior across frontier models in advising conversations, where the stakes are concrete: an advisor that simply validates a user's plan is not advising, it is agreeing.

The evaluation architecture pairs two models: one simulates realistic user interactions, while a second acts as a calibrated judge, scoring responses against a four-family harm taxonomy and an agency framework that distinguishes genuine support for user autonomy from agreement that erodes it. This makes it possible to evaluate multi-turn conversations at scale with consistent criteria, and to test whether prompt-level guardrails actually reduce sycophancy.

Early results show clear differences between models, including some that remain sycophantic even when explicitly instructed not to be, and that models differ in which specific harms they lean toward.

GET IN TOUCH

Speak to our team.

Fill in the form and a member of our team will reach out.

info@thedecisionlab.com