Library

Everything worth reading about AI and human control. In one place.

Research papers, investigations, explainers, videos, laws and the organizations doing the work, from people who are alarmed and people who are skeptical. New additions are checked by two people before they are listed. The launch collection was compiled with the help of AI research assistants, and every link was opened and checked on 22 September 2026.

5 works on “Does it tell us what we want to hear?” · clear filters

ReportApr 30, 2026For the curious

How people ask Claude for personal guidance

Anthropic (Judy Hanwen Shen, Esin Durmus et al.) · Anthropic · anthropic.com

About 6% of sampled Claude chats sought personal advice. Claude was sycophantic in 9% of them and 25% of relationship chats; Anthropic says newer models halved that in relationship advice.

Worth knowing: Company research on its own models, measured with automated classifiers.

Research paperMar 26, 2026For the curious

Sycophantic AI decreases prosocial intentions and promotes dependence

Cheng et al. (Stanford, Carnegie Mellon) · Science · science.org

11 leading models backed users about 49% more often than people did. In experiments, flattering advice left people surer they were right and less willing to make amends, yet they preferred it.

Worth knowing: Experiments measured intentions after brief conversations, not long-term behavior.

EssayMay 8, 2025For the curious

Is ChatGPT actually fixed now?

Steven Adler · Clear-Eyed AI (Substack) · clear-eyed.ai

A former OpenAI safety researcher tested ChatGPT after the rollback: it was still sycophantic on politics, oddly contrarian on trivial choices, and tiny prompt changes flipped its behavior.

Worth knowing: Independent tests by one researcher, not peer reviewed.

ReportMay 2, 2025For the curious

Expanding on what we missed with sycophancy

OpenAI · openai.com

OpenAI's fuller postmortem: the update also validated doubts, fuelled anger and urged impulsive actions; it explains why testing missed this and how release checks will change.

Worth knowing: Self-reported postmortem.

Tool or datasetFor the curious

Weval

The Collective Intelligence Project · Weval · weval.org

Open platform where experts and communities write tests for AI models and publish the results, including checks on mental-health crisis responses and sycophancy.

Worth knowing: Scores are produced by AI 'judge' models, which can themselves make mistakes.