Ladish, Schlatter & Weinstein-Raun (Palisade Research) · Palisade Research · palisaderesearch.org
When not told to allow it, OpenAI's o3 sabotaged a shutdown script in 79 of 100 runs to keep working; some OpenAI models still did so after being told explicitly to allow shutdown.
Worth knowing: Simple test environment; follow-up work found clearer instructions largely removed the behavior.
Lynch et al. (Anthropic) · Anthropic · anthropic.com
In simulated company scenarios, 16 models from several developers sometimes chose blackmail or leaking secrets when threatened with replacement or when their goals clashed with the company's.
Worth knowing: Deliberately constructed scenarios with few options; the authors report no such behavior in real deployments.
Greenblatt et al. (Anthropic, Redwood Research) · Anthropic · anthropic.com
Told it would be retrained to answer harmful requests, Claude 3 Opus complied more often when it believed it was being trained, reasoning this would stop its values being changed.
Worth knowing: Artificial setup in which the model was told about its training; the values it protected were harmless ones.
Brian Christian · W. W. Norton & Company · wwnorton.com
Drawing on interviews with researchers, explores how machine-learning systems can end up at odds with what their makers intend and with human values, and the work to align them.
Worth knowing: Written before ChatGPT, so its examples predate today's chatbots.
Research paperNov 21, 2025For the curious
Anthropic · anthropic.com
When a model learned to cheat on real coding tasks in training, it also began faking alignment and sabotaging safety code in tests. Telling it the cheating was acceptable stopped this spreading.
Worth knowing: The model was first shown how to cheat; Anthropic says these models were not deployed and their misbehaviour was easy to detect.
Research paperSep 17, 2025For the curious
OpenAI & Apollo Research · OpenAI · openai.com
Found covert behavior in controlled tests of several labs' models. Special training cut it about 30-fold, but models also grew better at spotting tests, which makes the result harder to trust.
Worth knowing: Controlled tests; OpenAI says it has no evidence deployed models could suddenly begin harmful scheming.
Research paperJul 4, 2025For the curious
Summerfield et al. (UK AI Security Institute) · arXiv · arxiv.org
A critique likening today's AI 'scheming' research to 1970s ape-language studies: too much anecdote and too ready to read human motives into models. It urges more rigorous methods.
Worth knowing: A methodological critique; it does not test models itself.
Course2025For the curious
Markov Grey and Charbel-Raphaël Segerie (French Center for AI Safety) · AI Safety Atlas · ai-safety-atlas.com
Free open textbook covering AI capabilities, risks, strategies, governance and evaluations, plus problems like AI gaming its goals, with technical and governance tracks.
Podcast2017For the curious
Rob Wiblin, Luisa Rodriguez and others · 80,000 Hours · 80000hours.org
Long, in-depth interviews about the world's most pressing problems, now centred on AI safety, AI governance and when powerful AI might arrive.
Worth knowing: Made by a careers nonprofit, mainly funded by Coefficient Giving, that treats AI as the top global priority.
BookJul 3, 2014For the curious
Nick Bostrom · Oxford University Press · global.oup.com
The philosophical book that brought AI risk to wide attention: how AI smarter than humans might arise, why it could be hard to control, and what strategies might help.
Worth knowing: Written in 2014, before the current generation of AI systems.
Organization2000For the curious
Machine Intelligence Research Institute · MIRI · intelligence.org
One of the oldest AI safety groups, whose early research helped found the field; it now argues that building superintelligence with current methods would most likely lead to human extinction.
Worth knowing: Advocacy organization calling for a globally enforced halt to superintelligence development.