Library

Everything worth reading about AI and human control. In one place.

Research papers, investigations, explainers, videos, laws and the organizations doing the work, from people who are alarmed and people who are skeptical. New additions are checked by two people before they are listed. The launch collection was compiled with the help of AI research assistants, and every link was opened and checked on 22 September 2026.

9 works on “Does it tell us what we want to hear?” · clear filters

EssentialIncidentApr 29, 2025For everyone

Sycophancy in GPT-4o: what happened and what we’re doing about it

OpenAI · openai.com

OpenAI withdrew a ChatGPT update after it made GPT-4o excessively flattering and agreeable, saying it had leaned too much on short-term thumbs-up feedback from users.

Worth knowing: The company's own account of a failure in its product.

EssentialResearch paperOct 20, 2023Technical

Towards Understanding Sycophancy in Language Models

Sharma et al. (Anthropic) · arXiv (ICLR 2024) · arxiv.org

Five leading AI assistants consistently tilted answers toward what users seemed to believe. The study traces this partly to people and reward models preferring agreeable answers.

Worth knowing: Lab study that includes the lab's own models.

ReportApr 30, 2026For the curious

How people ask Claude for personal guidance

Anthropic (Judy Hanwen Shen, Esin Durmus et al.) · Anthropic · anthropic.com

About 6% of sampled Claude chats sought personal advice. Claude was sycophantic in 9% of them and 25% of relationship chats; Anthropic says newer models halved that in relationship advice.

Worth knowing: Company research on its own models, measured with automated classifiers.

ArticleMar 26, 2026For everyone

AI overly affirms users asking for personal advice

Ula Chrobak · Stanford Report · news.stanford.edu

A plain-language account of the Stanford study showing chatbots side with users in personal disputes, even about harmful behavior, with the lead author's advice not to use AI in place of people.

Worth knowing: University news article about its own researchers' work.

Research paperMar 26, 2026For the curious

Sycophantic AI decreases prosocial intentions and promotes dependence

Cheng et al. (Stanford, Carnegie Mellon) · Science · science.org

11 leading models backed users about 49% more often than people did. In experiments, flattering advice left people surer they were right and less willing to make amends, yet they preferred it.

Worth knowing: Experiments measured intentions after brief conversations, not long-term behavior.

EssayMay 8, 2025For the curious

Is ChatGPT actually fixed now?

Steven Adler · Clear-Eyed AI (Substack) · clear-eyed.ai

A former OpenAI safety researcher tested ChatGPT after the rollback: it was still sycophantic on politics, oddly contrarian on trivial choices, and tiny prompt changes flipped its behavior.

Worth knowing: Independent tests by one researcher, not peer reviewed.

ReportMay 2, 2025For the curious

Expanding on what we missed with sycophancy

OpenAI · openai.com

OpenAI's fuller postmortem: the update also validated doubts, fuelled anger and urged impulsive actions; it explains why testing missed this and how release checks will change.

Worth knowing: Self-reported postmortem.

Research paperDec 19, 2022Technical

Discovering Language Model Behaviors with Model-Written Evaluations

Perez et al. (Anthropic) · arXiv · arxiv.org

Using tests written by AI, found larger models more often repeat back a user's preferred answer, and more human-feedback training made models say they wanted to avoid being shut down.

Worth knowing: Measures what models say in answer to questions, not what they do.

Tool or datasetFor the curious

Weval

The Collective Intelligence Project · Weval · weval.org

Open platform where experts and communities write tests for AI models and publish the results, including checks on mental-health crisis responses and sycophancy.

Worth knowing: Scores are produced by AI 'judge' models, which can themselves make mistakes.