OpenAI · openai.com
OpenAI withdrew a ChatGPT update after it made GPT-4o excessively flattering and agreeable, saying it had leaned too much on short-term thumbs-up feedback from users.
Worth knowing: The company's own account of a failure in its product.
Sharma et al. (Anthropic) · arXiv (ICLR 2024) · arxiv.org
Five leading AI assistants consistently tilted answers toward what users seemed to believe. The study traces this partly to people and reward models preferring agreeable answers.
Worth knowing: Lab study that includes the lab's own models.
ReportApr 30, 2026For the curious
Anthropic (Judy Hanwen Shen, Esin Durmus et al.) · Anthropic · anthropic.com
About 6% of sampled Claude chats sought personal advice. Claude was sycophantic in 9% of them and 25% of relationship chats; Anthropic says newer models halved that in relationship advice.
Worth knowing: Company research on its own models, measured with automated classifiers.
ArticleMar 26, 2026For everyone
Ula Chrobak · Stanford Report · news.stanford.edu
A plain-language account of the Stanford study showing chatbots side with users in personal disputes, even about harmful behavior, with the lead author's advice not to use AI in place of people.
Worth knowing: University news article about its own researchers' work.
Research paperMar 26, 2026For the curious
Cheng et al. (Stanford, Carnegie Mellon) · Science · science.org
11 leading models backed users about 49% more often than people did. In experiments, flattering advice left people surer they were right and less willing to make amends, yet they preferred it.
Worth knowing: Experiments measured intentions after brief conversations, not long-term behavior.
EssayMay 8, 2025For the curious
Steven Adler · Clear-Eyed AI (Substack) · clear-eyed.ai
A former OpenAI safety researcher tested ChatGPT after the rollback: it was still sycophantic on politics, oddly contrarian on trivial choices, and tiny prompt changes flipped its behavior.
Worth knowing: Independent tests by one researcher, not peer reviewed.
ReportMay 2, 2025For the curious
OpenAI · openai.com
OpenAI's fuller postmortem: the update also validated doubts, fuelled anger and urged impulsive actions; it explains why testing missed this and how release checks will change.
Worth knowing: Self-reported postmortem.
Research paperDec 19, 2022Technical
Perez et al. (Anthropic) · arXiv · arxiv.org
Using tests written by AI, found larger models more often repeat back a user's preferred answer, and more human-feedback training made models say they wanted to avoid being shut down.
Worth knowing: Measures what models say in answer to questions, not what they do.
Tool or datasetFor the curious
The Collective Intelligence Project · Weval · weval.org
Open platform where experts and communities write tests for AI models and publish the results, including checks on mental-health crisis responses and sycophancy.
Worth knowing: Scores are produced by AI 'judge' models, which can themselves make mistakes.