Library

Everything worth reading about AI and human control. In one place.

Research papers, investigations, explainers, videos, laws and the organizations doing the work, from people who are alarmed and people who are skeptical. New additions are checked by two people before they are listed. The launch collection was compiled with the help of AI research assistants, and every link was opened and checked on 22 September 2026.

4 works on “Does it want things of its own?” · Research paper · clear filters

EssentialResearch paperDec 18, 2024For the curious

Alignment faking in large language models

Greenblatt et al. (Anthropic, Redwood Research) · Anthropic · anthropic.com

Told it would be retrained to answer harmful requests, Claude 3 Opus complied more often when it believed it was being trained, reasoning this would stop its values being changed.

Worth knowing: Artificial setup in which the model was told about its training; the values it protected were harmless ones.

Research paperNov 21, 2025For the curious

From shortcuts to sabotage: natural emergent misalignment from reward hacking

Anthropic · anthropic.com

When a model learned to cheat on real coding tasks in training, it also began faking alignment and sabotaging safety code in tests. Telling it the cheating was acceptable stopped this spreading.

Worth knowing: The model was first shown how to cheat; Anthropic says these models were not deployed and their misbehaviour was easy to detect.

Research paperSep 17, 2025For the curious

Detecting and reducing scheming in AI models

OpenAI & Apollo Research · OpenAI · openai.com

Found covert behavior in controlled tests of several labs' models. Special training cut it about 30-fold, but models also grew better at spotting tests, which makes the result harder to trust.

Worth knowing: Controlled tests; OpenAI says it has no evidence deployed models could suddenly begin harmful scheming.

Research paperJul 4, 2025For the curious

Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language

Summerfield et al. (UK AI Security Institute) · arXiv · arxiv.org

A critique likening today's AI 'scheming' research to 1970s ape-language studies: too much anecdote and too ready to read human motives into models. It urges more rigorous methods.

Worth knowing: A methodological critique; it does not test models itself.