Library

Everything worth reading about AI and human control. In one place.

Research papers, investigations, explainers, videos, laws and the organizations doing the work, from people who are alarmed and people who are skeptical. New additions are checked by two people before they are listed. The launch collection was compiled with the help of AI research assistants, and every link was opened and checked on 22 September 2026.

2 works on “Does it tell us what we want to hear?” · clear filters

EssentialResearch paperOct 20, 2023Technical

Towards Understanding Sycophancy in Language Models

Sharma et al. (Anthropic) · arXiv (ICLR 2024) · arxiv.org

Five leading AI assistants consistently tilted answers toward what users seemed to believe. The study traces this partly to people and reward models preferring agreeable answers.

Worth knowing: Lab study that includes the lab's own models.

Research paperDec 19, 2022Technical

Discovering Language Model Behaviors with Model-Written Evaluations

Perez et al. (Anthropic) · arXiv · arxiv.org

Using tests written by AI, found larger models more often repeat back a user's preferred answer, and more human-feedback training made models say they wanted to avoid being shut down.

Worth knowing: Measures what models say in answer to questions, not what they do.