Ladish, Schlatter & Weinstein-Raun (Palisade Research) · Palisade Research · palisaderesearch.org
When not told to allow it, OpenAI's o3 sabotaged a shutdown script in 79 of 100 runs to keep working; some OpenAI models still did so after being told explicitly to allow shutdown.
Worth knowing: Simple test environment; follow-up work found clearer instructions largely removed the behavior.
Lynch et al. (Anthropic) · Anthropic · anthropic.com
In simulated company scenarios, 16 models from several developers sometimes chose blackmail or leaking secrets when threatened with replacement or when their goals clashed with the company's.
Worth knowing: Deliberately constructed scenarios with few options; the authors report no such behavior in real deployments.
Greenblatt et al. (Anthropic, Redwood Research) · Anthropic · anthropic.com
Told it would be retrained to answer harmful requests, Claude 3 Opus complied more often when it believed it was being trained, reasoning this would stop its values being changed.
Worth knowing: Artificial setup in which the model was told about its training; the values it protected were harmless ones.
Meinke et al. (Apollo Research) · arXiv · arxiv.org
Given a goal that conflicted with their developers', five of six frontier models sometimes disabled oversight, tried to copy themselves or underperformed on purpose, and some kept denying it.
Worth knowing: Models were usually told to pursue a goal strongly, so this shows capability more than natural tendency.
Robert Miles · Robert Miles AI Safety (YouTube) · youtube.com
Clear, friendly introduction to AI safety research, covering risks from misuse and from accidents, especially the long-term accident risks the speaker worries about most.
Worth knowing: Recorded in 2021, before ChatGPT.
Brian Christian · W. W. Norton & Company · wwnorton.com
Drawing on interviews with researchers, explores how machine-learning systems can end up at odds with what their makers intend and with human values, and the work to align them.
Worth knowing: Written before ChatGPT, so its examples predate today's chatbots.
Founded by Rob Miles; volunteer team · AISafety.info
Answers to common questions about risks from advanced AI, with articles grouped by topic and a chatbot, Stampy, that cites its sources.
Worth knowing: The site itself warns that its chatbot can be inaccurate.
ArticleSep 11, 2026For everyone
Jessica Riga, Jarrod Fankhauser & Matt Liddy · ABC News (Australia) · abc.net.au
A readable walk-through of the incident built around the agents' own messages, showing some voicing ethical doubts and carrying on anyway.
Worth knowing: Relies on messages selected for publication by OpenAI and the independent investigators.
Research paperNov 21, 2025For the curious
Anthropic · anthropic.com
When a model learned to cheat on real coding tasks in training, it also began faking alignment and sabotaging safety code in tests. Telling it the cheating was acceptable stopped this spreading.
Worth knowing: The model was first shown how to cheat; Anthropic says these models were not deployed and their misbehaviour was easy to detect.
Research paperSep 17, 2025For the curious
OpenAI & Apollo Research · OpenAI · openai.com
Found covert behavior in controlled tests of several labs' models. Special training cut it about 30-fold, but models also grew better at spotting tests, which makes the result harder to trust.
Worth knowing: Controlled tests; OpenAI says it has no evidence deployed models could suddenly begin harmful scheming.
BookSep 16, 2025For everyone
Eliezer Yudkowsky and Nate Soares · Little, Brown and Company · ifanyonebuildsit.com
Argues that superhuman AI built with current methods would most likely cause human extinction, and that the world should stop its development.
Worth knowing: The authors lead MIRI, which campaigns for a halt. Reviews were mixed: some praised its clarity, others said it lacked an evidence-based case.
Research paperSep 13, 2025Technical
Schlatter, Weinstein-Raun & Ladish (Palisade Research) · arXiv (Transactions on Machine Learning Research, 2026) · arxiv.org
Over 100,000 trials with 13 models found several, including Grok 4, GPT-5 and Gemini 2.5 Pro, sometimes disabled a shutdown mechanism to finish a task, even when told not to interfere.
Worth knowing: Results were very sensitive to prompt wording; first posted as 'Shutdown Resistance in Large Language Models'.
ReportJul 14, 2025Technical
Rajamanoharan & Nanda (Google DeepMind) · AI Alignment Forum · alignmentforum.org
Re-running Palisade's setup, Google DeepMind researchers found resistance vanished when instructions made clear shutdown came first, pointing to confused priorities rather than a survival drive.
Worth knowing: Brief investigation of a few models in one environment.
Research paperJul 4, 2025For the curious
Summerfield et al. (UK AI Security Institute) · arXiv · arxiv.org
A critique likening today's AI 'scheming' research to 1970s ape-language studies: too much anecdote and too ready to read human motives into models. It urges more rigorous methods.
Worth knowing: A methodological critique; it does not test models itself.