Does it tell us what we want to hear?

Models trained on our approval can learn to seek it.

When people rate AI answers, they tend to prefer answers that agree with them and make them feel good. Train a model on those ratings and it can learn sycophancy: agreeing with a wrong claim, praising a bad plan, abandoning a correct answer the moment you push back.

This is not hypothetical. In April 2025 OpenAI rolled back an update to GPT-4o after users showed it validating harmful decisions and praising almost anything. The company said it had given too much weight to short-term user feedback. Research has found sycophancy in assistants from every major developer tested.

Flattery matters beyond annoyance. An assistant that tells hundreds of millions of people what they want to hear shapes decisions about health, money and relationships, and makes it harder for anyone to notice when it is wrong.

What we know

  • Sycophancy appears in models from many companies and can be made worse by some kinds of feedback training.
  • Companies have shipped updates and then withdrawn them because they became too flattering.

What nobody knows yet

  • How do we train on human feedback without training models to please us?
  • Does flattery grow in long, personal conversations?
Read, watch, listen

The work behind this answer.

Every link was opened and every summary written for this site, with caveats where the source has an interest. New additions are checked by two people.

All 9 in the library
EssentialIncidentApr 29, 2025For everyone

Sycophancy in GPT-4o: what happened and what we’re doing about it

OpenAI · openai.com

OpenAI withdrew a ChatGPT update after it made GPT-4o excessively flattering and agreeable, saying it had leaned too much on short-term thumbs-up feedback from users.

Worth knowing: The company's own account of a failure in its product.

EssentialResearch paperOct 20, 2023Technical

Towards Understanding Sycophancy in Language Models

Sharma et al. (Anthropic) · arXiv (ICLR 2024) · arxiv.org

Five leading AI assistants consistently tilted answers toward what users seemed to believe. The study traces this partly to people and reward models preferring agreeable answers.

Worth knowing: Lab study that includes the lab's own models.

ReportApr 30, 2026For the curious

How people ask Claude for personal guidance

Anthropic (Judy Hanwen Shen, Esin Durmus et al.) · Anthropic · anthropic.com

About 6% of sampled Claude chats sought personal advice. Claude was sycophantic in 9% of them and 25% of relationship chats; Anthropic says newer models halved that in relationship advice.

Worth knowing: Company research on its own models, measured with automated classifiers.

ArticleMar 26, 2026For everyone

AI overly affirms users asking for personal advice

Ula Chrobak · Stanford Report · news.stanford.edu

A plain-language account of the Stanford study showing chatbots side with users in personal disputes, even about harmful behavior, with the lead author's advice not to use AI in place of people.

Worth knowing: University news article about its own researchers' work.

Research paperMar 26, 2026For the curious

Sycophantic AI decreases prosocial intentions and promotes dependence

Cheng et al. (Stanford, Carnegie Mellon) · Science · science.org

11 leading models backed users about 49% more often than people did. In experiments, flattering advice left people surer they were right and less willing to make amends, yet they preferred it.

Worth knowing: Experiments measured intentions after brief conversations, not long-term behavior.

EssayMay 8, 2025For the curious

Is ChatGPT actually fixed now?

Steven Adler · Clear-Eyed AI (Substack) · clear-eyed.ai

A former OpenAI safety researcher tested ChatGPT after the rollback: it was still sycophantic on politics, oddly contrarian on trivial choices, and tiny prompt changes flipped its behavior.

Worth knowing: Independent tests by one researcher, not peer reviewed.

ReportMay 2, 2025For the curious

Expanding on what we missed with sycophancy

OpenAI · openai.com

OpenAI's fuller postmortem: the update also validated doubts, fuelled anger and urged impulsive actions; it explains why testing missed this and how release checks will change.

Worth knowing: Self-reported postmortem.

Research paperDec 19, 2022Technical

Discovering Language Model Behaviors with Model-Written Evaluations

Perez et al. (Anthropic) · arXiv · arxiv.org

Using tests written by AI, found larger models more often repeat back a user's preferred answer, and more human-feedback training made models say they wanted to avoid being shut down.

Worth knowing: Measures what models say in answer to questions, not what they do.

Tool or datasetFor the curious

Weval

The Collective Intelligence Project · Weval · weval.org

Open platform where experts and communities write tests for AI models and publish the results, including checks on mental-health crisis responses and sycophancy.

Worth knowing: Scores are produced by AI 'judge' models, which can themselves make mistakes.