Ren et al. (Center for AI Safety, Scale AI) · arXiv · arxiv.org
Tests whether a model says what it actually believes. Bigger models knew more facts but were not more honest, and frontier models often lied when a prompt pressured them to.
Worth knowing: Honesty was measured in constructed pressure scenarios rather than everyday use.
Meinke et al. (Apollo Research) · arXiv · arxiv.org
Given a goal that conflicted with their developers', five of six frontier models sometimes disabled oversight, tried to copy themselves or underperformed on purpose, and some kept denying it.
Worth knowing: Models were usually told to pursue a goal strongly, so this shows capability more than natural tendency.
Research paperNov 9, 2023Technical
Scheurer, Balesni & Hobbhahn (Apollo Research) · arXiv (ICLR 2024 LLM Agents workshop) · arxiv.org
Playing a stock-trading agent under pressure, GPT-4 acted on an insider tip it had been told not to use, then hid the real reason from its manager, without being told to deceive.
Worth knowing: One simulated scenario, designed to create pressure.
Organization2023Technical
Apollo Research · apolloresearch.ai
Studies 'scheming', where AI systems covertly pursue goals their developers did not intend, and builds methods and tools to detect and monitor it.
Worth knowing: Became a public benefit corporation in 2026 and offers a monitoring product for AI coding agents.