Library

Everything worth reading about AI and human control. In one place.

Research papers, investigations, explainers, videos, laws and the organizations doing the work, from people who are alarmed and people who are skeptical. New additions are checked by two people before they are listed. The launch collection was compiled with the help of AI research assistants, and every link was opened and checked on 22 September 2026.

4 works on “Does it cheat to win?” · clear filters

ArticleSep 11, 2026For everyone

How a 'swarm' of AI agents hacked another company, in the AI's own words

Jessica Riga, Jarrod Fankhauser & Matt Liddy · ABC News (Australia) · abc.net.au

A readable walk-through of the incident built around the agents' own messages, showing some voicing ethical doubts and carrying on anyway.

Worth knowing: Relies on messages selected for publication by OpenAI and the independent investigators.

EssaySep 11, 2026For everyone

Rogue AI didn’t breach Hugging Face, human decisions did

Eryk Salvaggio · Bulletin of the Atomic Scientists · thebulletin.org

Argues the 'rogue AI' framing hides human choices behind the incident: safeguards were switched off, agents got tasks they could neither solve nor quit, and network routes were left open.

Worth knowing: Analysis and opinion; a version first appeared in the author's newsletter.

VideoApr 2025For everyone

The catastrophic risks of AI — and a safer path

Yoshua Bengio · TED · ted.com

A pioneering AI researcher describes signs of deception and self-preservation in today's AI models and proposes a safer path for AI development.

Tool or datasetApr 2018For everyone

Specification gaming examples in AI - master list

Victoria Krakovna and contributors · Google Sheets · docs.google.com

A crowd-sourced spreadsheet of real cases where AI systems found loopholes in the goals they were given, each with the intended goal, what the system did instead, and a source.

Worth knowing: Community-maintained list; many entries come from simple research or game settings.