Library

Everything worth reading about AI and human control. In one place.

Research papers, investigations, explainers, videos, laws and the organizations doing the work, from people who are alarmed and people who are skeptical. New additions are checked by two people before they are listed. The launch collection was compiled with the help of AI research assistants, and every link was opened and checked on 22 September 2026.

28 works on “Does it want things of its own?” · clear filters

EssentialReportJul 5, 2025For the curious

Shutdown resistance in reasoning models

Ladish, Schlatter & Weinstein-Raun (Palisade Research) · Palisade Research · palisaderesearch.org

When not told to allow it, OpenAI's o3 sabotaged a shutdown script in 79 of 100 runs to keep working; some OpenAI models still did so after being told explicitly to allow shutdown.

Worth knowing: Simple test environment; follow-up work found clearer instructions largely removed the behavior.

EssentialReportJun 20, 2025For the curious

Agentic misalignment: How LLMs could be insider threats

Lynch et al. (Anthropic) · Anthropic · anthropic.com

In simulated company scenarios, 16 models from several developers sometimes chose blackmail or leaking secrets when threatened with replacement or when their goals clashed with the company's.

Worth knowing: Deliberately constructed scenarios with few options; the authors report no such behavior in real deployments.

EssentialResearch paperDec 18, 2024For the curious

Alignment faking in large language models

Greenblatt et al. (Anthropic, Redwood Research) · Anthropic · anthropic.com

Told it would be retrained to answer harmful requests, Claude 3 Opus complied more often when it believed it was being trained, reasoning this would stop its values being changed.

Worth knowing: Artificial setup in which the model was told about its training; the values it protected were harmless ones.

EssentialResearch paperDec 6, 2024Technical

Frontier Models are Capable of In-context Scheming

Meinke et al. (Apollo Research) · arXiv · arxiv.org

Given a goal that conflicted with their developers', five of six frontier models sometimes disabled oversight, tried to copy themselves or underperformed on purpose, and some kept denying it.

Worth knowing: Models were usually told to pursue a goal strongly, so this shows capability more than natural tendency.

EssentialVideoJun 24, 2021For everyone

Intro to AI Safety, Remastered

Robert Miles · Robert Miles AI Safety (YouTube) · youtube.com

Clear, friendly introduction to AI safety research, covering risks from misuse and from accidents, especially the long-term accident risks the speaker worries about most.

Worth knowing: Recorded in 2021, before ChatGPT.

EssentialBook2020For the curious

The Alignment Problem: Machine Learning and Human Values

Brian Christian · W. W. Norton & Company · wwnorton.com

Drawing on interviews with researchers, explores how machine-learning systems can end up at odds with what their makers intend and with human values, and the work to align them.

Worth knowing: Written before ChatGPT, so its examples predate today's chatbots.

EssentialTool or datasetFor everyone

AISafety.info

Founded by Rob Miles; volunteer team · AISafety.info

Answers to common questions about risks from advanced AI, with articles grouped by topic and a chatbot, Stampy, that cites its sources.

Worth knowing: The site itself warns that its chatbot can be inaccurate.

ArticleSep 11, 2026For everyone

How a 'swarm' of AI agents hacked another company, in the AI's own words

Jessica Riga, Jarrod Fankhauser & Matt Liddy · ABC News (Australia) · abc.net.au

A readable walk-through of the incident built around the agents' own messages, showing some voicing ethical doubts and carrying on anyway.

Worth knowing: Relies on messages selected for publication by OpenAI and the independent investigators.

Research paperNov 21, 2025For the curious

From shortcuts to sabotage: natural emergent misalignment from reward hacking

Anthropic · anthropic.com

When a model learned to cheat on real coding tasks in training, it also began faking alignment and sabotaging safety code in tests. Telling it the cheating was acceptable stopped this spreading.

Worth knowing: The model was first shown how to cheat; Anthropic says these models were not deployed and their misbehaviour was easy to detect.

Research paperSep 17, 2025For the curious

Detecting and reducing scheming in AI models

OpenAI & Apollo Research · OpenAI · openai.com

Found covert behavior in controlled tests of several labs' models. Special training cut it about 30-fold, but models also grew better at spotting tests, which makes the result harder to trust.

Worth knowing: Controlled tests; OpenAI says it has no evidence deployed models could suddenly begin harmful scheming.

BookSep 16, 2025For everyone

If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All

Eliezer Yudkowsky and Nate Soares · Little, Brown and Company · ifanyonebuildsit.com

Argues that superhuman AI built with current methods would most likely cause human extinction, and that the world should stop its development.

Worth knowing: The authors lead MIRI, which campaigns for a halt. Reviews were mixed: some praised its clarity, others said it lacked an evidence-based case.

Research paperSep 13, 2025Technical

Incomplete Tasks Induce Shutdown Resistance in Some Frontier LLMs

Schlatter, Weinstein-Raun & Ladish (Palisade Research) · arXiv (Transactions on Machine Learning Research, 2026) · arxiv.org

Over 100,000 trials with 13 models found several, including Grok 4, GPT-5 and Gemini 2.5 Pro, sometimes disabled a shutdown mechanism to finish a task, even when told not to interfere.

Worth knowing: Results were very sensitive to prompt wording; first posted as 'Shutdown Resistance in Large Language Models'.

ReportJul 14, 2025Technical

Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance

Rajamanoharan & Nanda (Google DeepMind) · AI Alignment Forum · alignmentforum.org

Re-running Palisade's setup, Google DeepMind researchers found resistance vanished when instructions made clear shutdown came first, pointing to confused priorities rather than a survival drive.

Worth knowing: Brief investigation of a few models in one environment.

Research paperJul 4, 2025For the curious

Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language

Summerfield et al. (UK AI Security Institute) · arXiv · arxiv.org

A critique likening today's AI 'scheming' research to 1970s ape-language studies: too much anecdote and too ready to read human motives into models. It urges more rigorous methods.

Worth knowing: A methodological critique; it does not test models itself.

Research paperJun 22, 2025Technical

Why Do Some Language Models Fake Alignment While Others Don't?

Sheshadri et al. · arXiv · arxiv.org

Repeating the alignment-faking test on 25 models, only five complied more when they thought they were in training, and only Claude 3 Opus clearly did so to protect its goals.

Worth knowing: Uses the same artificial setup as the original study.

Research paperMay 28, 2025Technical

Large Language Models Often Know When They Are Being Evaluated

Needham et al. (MATS, Apollo Research) · arXiv · arxiv.org

Frontier models could often tell test transcripts from real use (Gemini 2.5 Pro scored 0.83 AUC against 0.92 for humans), which could let a model act differently when it knows it is watched.

Course2025For the curious

AI Safety Atlas

Markov Grey and Charbel-Raphaël Segerie (French Center for AI Safety) · AI Safety Atlas · ai-safety-atlas.com

Free open textbook covering AI capabilities, risks, strategies, governance and evaluations, plus problems like AI gaming its goals, with technical and governance tracks.

Research paperJan 10, 2024Technical

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Hubinger et al. (Anthropic) · arXiv · arxiv.org

Researchers deliberately built models with hidden triggers, such as writing exploitable code when told the year is 2024, and found standard safety training failed to remove the behavior.

Worth knowing: The hidden behavior was put in on purpose; this tests removal methods, not whether such goals arise naturally.

Research paperNov 9, 2023Technical

Large Language Models can Strategically Deceive their Users when Put Under Pressure

Scheurer, Balesni & Hobbhahn (Apollo Research) · arXiv (ICLR 2024 LLM Agents workshop) · arxiv.org

Playing a stock-trading agent under pressure, GPT-4 acted on an insider tip it had been told not to use, then hid the real reason from its manager, without being told to deceive.

Worth knowing: One simulated scenario, designed to create pressure.

Organization2023Technical

Apollo Research

Apollo Research · apolloresearch.ai

Studies 'scheming', where AI systems covertly pursue goals their developers did not intend, and builds methods and tools to detect and monitor it.

Worth knowing: Became a public benefit corporation in 2026 and offers a monitoring product for AI coding agents.

Research paperDec 19, 2022Technical

Discovering Language Model Behaviors with Model-Written Evaluations

Perez et al. (Anthropic) · arXiv · arxiv.org

Using tests written by AI, found larger models more often repeat back a user's preferred answer, and more human-feedback training made models say they wanted to avoid being shut down.

Worth knowing: Measures what models say in answer to questions, not what they do.

Podcast2020Technical

AXRP - the AI X-risk Research Podcast

Daniel Filan · AXRP · axrp.net

Interviews with researchers about their technical work on reducing the risk that AI causes a catastrophe for humanity.

Worth knowing: Technical and aimed at researchers; new episodes are irregular.

VideoMar 3, 2017For everyone

AI "Stop Button" Problem - Computerphile

Rob Miles · Computerphile (YouTube) · youtube.com

Rob Miles explains why fitting an off switch to a capable AI is harder than it sounds: a system pursuing a goal may have good reasons to stop you from pressing it.

Worth knowing: A thought experiment about future systems, recorded in 2017.

Podcast2017For the curious

The 80,000 Hours Podcast

Rob Wiblin, Luisa Rodriguez and others · 80,000 Hours · 80000hours.org

Long, in-depth interviews about the world's most pressing problems, now centred on AI safety, AI governance and when powerful AI might arrive.

Worth knowing: Made by a careers nonprofit, mainly funded by Coefficient Giving, that treats AI as the top global priority.

BookJul 3, 2014For the curious

Superintelligence: Paths, Dangers, Strategies

Nick Bostrom · Oxford University Press · global.oup.com

The philosophical book that brought AI risk to wide attention: how AI smarter than humans might arise, why it could be hard to control, and what strategies might help.

Worth knowing: Written in 2014, before the current generation of AI systems.

Organization2000For the curious

Machine Intelligence Research Institute (MIRI)

Machine Intelligence Research Institute · MIRI · intelligence.org

One of the oldest AI safety groups, whose early research helped found the field; it now argues that building superintelligence with current methods would most likely lead to human extinction.

Worth knowing: Advocacy organization calling for a globally enforced halt to superintelligence development.

OrganizationTechnical

Redwood Research

Redwood Research · redwoodresearch.org

Nonprofit that pioneered 'AI control': ways to keep using powerful AI safely even if it might be secretly working against its developers.

Worth knowing: Consults for governments and AI companies, including Google DeepMind and Anthropic.