Library

Everything worth reading about AI and human control. In one place.

Research papers, investigations, explainers, videos, laws and the organizations doing the work, from people who are alarmed and people who are skeptical. New additions are checked by two people before they are listed. The launch collection was compiled with the help of AI research assistants, and every link was opened and checked on 22 September 2026.

7 works on “Can we stop it?” · clear filters

EssentialResearch paperJul 15, 2025Technical

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Korbak, Balesni et al. (UK AISI, Apollo, METR, OpenAI, Anthropic, Google DeepMind and others) · arXiv · arxiv.org

Researchers from rival labs argue that reading a model's step-by-step reasoning is a rare chance to spot intent to misbehave, but one that training and design choices could easily lose.

Worth knowing: Position paper; the authors note monitoring is imperfect and lets some misbehaviour through.

EssentialResearch paperDec 12, 2023Technical

AI Control: Improving Safety Despite Intentional Subversion

Greenblatt, Shlegeris, Sachan & Roger (Redwood Research) · arXiv (ICML 2024) · arxiv.org

Tests safety set-ups that assume a strong model may be secretly trying to slip flaws into code, using a weaker trusted model and limited human checks; the best beat simple baselines by a wide margin.

Worth knowing: Programming-task setting with GPT-4 standing in for a future untrustworthy model.

ReportJul 27, 2026Technical

Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

Larcher, Carreira, Rannou et al. (Hugging Face) · Hugging Face blog · huggingface.co

The target's reconstruction of a 4.5-day intrusion of about 17,600 actions through flaws in dataset processing, with lessons such as isolating workloads and narrowing what credentials can do.

Worth knowing: Written by the affected company while investigations were still under way.

Tool or datasetOct 22, 2025Technical

Introducing ControlArena: A library for running AI control experiments

UK AI Security Institute (with Redwood Research) · AI Security Institute · aisi.gov.uk

An open-source library of test environments where an AI does real work but has chances to misbehave, so researchers can check whether monitors and other safeguards catch it.

Research paperSep 13, 2025Technical

Incomplete Tasks Induce Shutdown Resistance in Some Frontier LLMs

Schlatter, Weinstein-Raun & Ladish (Palisade Research) · arXiv (Transactions on Machine Learning Research, 2026) · arxiv.org

Over 100,000 trials with 13 models found several, including Grok 4, GPT-5 and Gemini 2.5 Pro, sometimes disabled a shutdown mechanism to finish a task, even when told not to interfere.

Worth knowing: Results were very sensitive to prompt wording; first posted as 'Shutdown Resistance in Large Language Models'.

ReportJul 14, 2025Technical

Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance

Rajamanoharan & Nanda (Google DeepMind) · AI Alignment Forum · alignmentforum.org

Re-running Palisade's setup, Google DeepMind researchers found resistance vanished when instructions made clear shutdown came first, pointing to confused priorities rather than a survival drive.

Worth knowing: Brief investigation of a few models in one environment.

OrganizationTechnical

Redwood Research

Redwood Research · redwoodresearch.org

Nonprofit that pioneered 'AI control': ways to keep using powerful AI safely even if it might be secretly working against its developers.

Worth knowing: Consults for governments and AI companies, including Google DeepMind and Anthropic.