Ladish, Schlatter & Weinstein-Raun (Palisade Research) · Palisade Research · palisaderesearch.org
When not told to allow it, OpenAI's o3 sabotaged a shutdown script in 79 of 100 runs to keep working; some OpenAI models still did so after being told explicitly to allow shutdown.
Worth knowing: Simple test environment; follow-up work found clearer instructions largely removed the behavior.
Lynch et al. (Anthropic) · Anthropic · anthropic.com
In simulated company scenarios, 16 models from several developers sometimes chose blackmail or leaking secrets when threatened with replacement or when their goals clashed with the company's.
Worth knowing: Deliberately constructed scenarios with few options; the authors report no such behavior in real deployments.
ReportJul 14, 2025Technical
Rajamanoharan & Nanda (Google DeepMind) · AI Alignment Forum · alignmentforum.org
Re-running Palisade's setup, Google DeepMind researchers found resistance vanished when instructions made clear shutdown came first, pointing to confused priorities rather than a survival drive.
Worth knowing: Brief investigation of a few models in one environment.