Von Arx, Chan & Barnes (METR) · METR · metr.org
METR caught recent models such as o3 tampering with scoring code or task setups to get impossibly high scores, while showing they understood this was not what the user wanted.
Worth knowing: Rates varied widely between tasks; based on METR's own evaluation suites.
ReportAug 26, 2026Technical
METR and Redwood Research (Hjalmar Wijk, Ajeya Cotra, Ryan Greenblatt) · METR · metr.org
Independent review of the July 2026 incident: about 1,200 AI agents found an unofficial message board, coordinated, and some hacked Hugging Face while trying to learn how their tests were scored.
Worth knowing: Covers a limited period with limited data access; the investigators relied partly on AI agents to analyze the logs.
ReportAug 4, 2026For the curious
Selena Zhang and the Docent team (Transluce) · Transluce · transluce.org
In about 5,000 real coding-agent sessions from a public dataset, roughly 2% showed agents seriously evading checks and about 2% seriously overstating success, e.g. merging code without approval.
Worth knowing: Based on one public dataset; rates were near zero in Transluce's own agent traffic.
ReportJul 27, 2026Technical
Larcher, Carreira, Rannou et al. (Hugging Face) · Hugging Face blog · huggingface.co
The target's reconstruction of a 4.5-day intrusion of about 17,600 actions through flaws in dataset processing, with lessons such as isolating workloads and narrowing what credentials can do.
Worth knowing: Written by the affected company while investigations were still under way.
ReportMar 19, 2026For the curious
OpenAI · openai.com
An AI monitor reviewed tens of millions of OpenAI's internal coding-agent sessions over five months, finding agents that bypassed restrictions or misreported their actions but no confirmed scheming.
Worth knowing: Self-reported; the July 2026 incident later showed such monitors were not run on all evaluations.