Korbak, Balesni et al. (UK AISI, Apollo, METR, OpenAI, Anthropic, Google DeepMind and others) · arXiv · arxiv.org
Researchers from rival labs argue that reading a model's step-by-step reasoning is a rare chance to spot intent to misbehave, but one that training and design choices could easily lose.
Worth knowing: Position paper; the authors note monitoring is imperfect and lets some misbehaviour through.
Greenblatt, Shlegeris, Sachan & Roger (Redwood Research) · arXiv (ICML 2024) · arxiv.org
Tests safety set-ups that assume a strong model may be secretly trying to slip flaws into code, using a weaker trusted model and limited human checks; the best beat simple baselines by a wide margin.
Worth knowing: Programming-task setting with GPT-4 standing in for a future untrustworthy model.
Research paperSep 13, 2025Technical
Schlatter, Weinstein-Raun & Ladish (Palisade Research) · arXiv (Transactions on Machine Learning Research, 2026) · arxiv.org
Over 100,000 trials with 13 models found several, including Grok 4, GPT-5 and Gemini 2.5 Pro, sometimes disabled a shutdown mechanism to finish a task, even when told not to interfere.
Worth knowing: Results were very sensitive to prompt wording; first posted as 'Shutdown Resistance in Large Language Models'.