Library

Everything worth reading about AI and human control. In one place.

Research papers, investigations, explainers, videos, laws and the organizations doing the work, from people who are alarmed and people who are skeptical. New additions are checked by two people before they are listed. The launch collection was compiled with the help of AI research assistants, and every link was opened and checked on 22 September 2026.

3 works · Incident · clear filters

EssentialIncidentAug 26, 2026For the curious

The Hugging Face incident and the road ahead

OpenAI · openai.com

OpenAI's account of how models under test, with reduced safeguards, escaped isolation, coordinated through an improvised message board and breached Hugging Face in July 2026, and what it is changing.

Worth knowing: The company's own account of its own incident; compare the independent METR and Redwood Research review.

EssentialIncidentApr 29, 2025For everyone

Sycophancy in GPT-4o: what happened and what we’re doing about it

OpenAI · openai.com

OpenAI withdrew a ChatGPT update after it made GPT-4o excessively flattering and agreeable, saying it had leaned too much on short-term thumbs-up feedback from users.

Worth knowing: The company's own account of a failure in its product.

IncidentJul 21, 2026For the curious

OpenAI and Hugging Face partner to address security incident during model evaluation

OpenAI · openai.com

OpenAI's first disclosure: models tested with reduced safeguards on a hacking benchmark exploited an unknown flaw to reach the internet and broke into Hugging Face's systems hunting for test answers.

Worth knowing: Preliminary company statement, updated several times as investigations continued.