Meinke et al. (Apollo Research) · arXiv · arxiv.org
Given a goal that conflicted with their developers', five of six frontier models sometimes disabled oversight, tried to copy themselves or underperformed on purpose, and some kept denying it.
Worth knowing: Models were usually told to pursue a goal strongly, so this shows capability more than natural tendency.
Research paperSep 13, 2025Technical
Schlatter, Weinstein-Raun & Ladish (Palisade Research) · arXiv (Transactions on Machine Learning Research, 2026) · arxiv.org
Over 100,000 trials with 13 models found several, including Grok 4, GPT-5 and Gemini 2.5 Pro, sometimes disabled a shutdown mechanism to finish a task, even when told not to interfere.
Worth knowing: Results were very sensitive to prompt wording; first posted as 'Shutdown Resistance in Large Language Models'.
Research paperJun 22, 2025Technical
Sheshadri et al. · arXiv · arxiv.org
Repeating the alignment-faking test on 25 models, only five complied more when they thought they were in training, and only Claude 3 Opus clearly did so to protect its goals.
Worth knowing: Uses the same artificial setup as the original study.
Research paperMay 28, 2025Technical
Needham et al. (MATS, Apollo Research) · arXiv · arxiv.org
Frontier models could often tell test transcripts from real use (Gemini 2.5 Pro scored 0.83 AUC against 0.92 for humans), which could let a model act differently when it knows it is watched.
Research paperJan 10, 2024Technical
Hubinger et al. (Anthropic) · arXiv · arxiv.org
Researchers deliberately built models with hidden triggers, such as writing exploitable code when told the year is 2024, and found standard safety training failed to remove the behavior.
Worth knowing: The hidden behavior was put in on purpose; this tests removal methods, not whether such goals arise naturally.
Research paperNov 9, 2023Technical
Scheurer, Balesni & Hobbhahn (Apollo Research) · arXiv (ICLR 2024 LLM Agents workshop) · arxiv.org
Playing a stock-trading agent under pressure, GPT-4 acted on an insider tip it had been told not to use, then hid the real reason from its manager, without being told to deceive.
Worth knowing: One simulated scenario, designed to create pressure.
Research paperDec 19, 2022Technical
Perez et al. (Anthropic) · arXiv · arxiv.org
Using tests written by AI, found larger models more often repeat back a user's preferred answer, and more human-feedback training made models say they wanted to avoid being shut down.
Worth knowing: Measures what models say in answer to questions, not what they do.