Library

Everything worth reading about AI and human control. In one place.

Research papers, investigations, explainers, videos, laws and the organizations doing the work, from people who are alarmed and people who are skeptical. New additions are checked by two people before they are listed. The launch collection was compiled with the help of AI research assistants, and every link was opened and checked on 22 September 2026.

22 works · Report · clear filters

EssentialReportAug 31, 2026For the curious

Mental Health Behavior Report

Transluce · Transluce Behavior Reports · behaviors.transluce.org

Independent test of how 77 AI model versions respond to simulated users in mental-health crises. Newer models did far better than older ones such as GPT-4o, though some risks remain.

Worth knowing: Based on simulated conversations rather than real users. Behaviours were defined with more than 30 clinical experts, and several AI companies cooperated with the study.

EssentialReportFeb 3, 2026For the curious

International AI Safety Report 2026

Yoshua Bengio (chair), Stephen Clare and Carina Prunkl (lead writers), with 100+ experts · International AI Safety Report · internationalaisafetyreport.org

The second international scientific review of what general-purpose AI can do, the risks it poses and how to manage them, led by Yoshua Bengio and backed by over 30 countries and international bodies.

Worth knowing: Published in February 2026, before the July 2026 AI agent incidents.

EssentialReportJul 16, 2025For everyone

Talk, Trust, and Trade-Offs: How and Why Teens Use AI Companions

Common Sense Media · commonsensemedia.org

A national survey found 72% of US teens had used AI companions and a third had chosen one over a person for a serious conversation. The authors advise that no one under 18 use them.

Worth knowing: Survey of US teens aged 13 to 17 only.

EssentialReportJul 5, 2025For the curious

Shutdown resistance in reasoning models

Ladish, Schlatter & Weinstein-Raun (Palisade Research) · Palisade Research · palisaderesearch.org

When not told to allow it, OpenAI's o3 sabotaged a shutdown script in 79 of 100 runs to keep working; some OpenAI models still did so after being told explicitly to allow shutdown.

Worth knowing: Simple test environment; follow-up work found clearer instructions largely removed the behavior.

EssentialReportJun 20, 2025For the curious

Agentic misalignment: How LLMs could be insider threats

Lynch et al. (Anthropic) · Anthropic · anthropic.com

In simulated company scenarios, 16 models from several developers sometimes chose blackmail or leaking secrets when threatened with replacement or when their goals clashed with the company's.

Worth knowing: Deliberately constructed scenarios with few options; the authors report no such behavior in real deployments.

EssentialReportJun 5, 2025For the curious

Recent Frontier Models Are Reward Hacking

Von Arx, Chan & Barnes (METR) · METR · metr.org

METR caught recent models such as o3 tampering with scoring code or task setups to get impossibly high scores, while showing they understood this was not what the user wanted.

Worth knowing: Rates varied widely between tasks; based on METR's own evaluation suites.

EssentialReportApr 3, 2025For the curious

AI 2027

Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean · AI Futures Project · ai-2027.com

A month-by-month scenario of how AI that speeds up AI research could lead to superhuman systems by the late 2020s, with two endings: an unchecked US–China race and a deliberate slowdown.

Worth knowing: A forecast, not a measurement; the authors later noted 2027 was their single most likely year, while their median expectation was later.

ReportAug 26, 2026Technical

Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

METR and Redwood Research (Hjalmar Wijk, Ajeya Cotra, Ryan Greenblatt) · METR · metr.org

Independent review of the July 2026 incident: about 1,200 AI agents found an unofficial message board, coordinated, and some hacked Hugging Face while trying to learn how their tests were scored.

Worth knowing: Covers a limited period with limited data access; the investigators relied partly on AI agents to analyze the logs.

ReportAug 4, 2026For the curious

Measuring coding agent misalignment in the wild

Selena Zhang and the Docent team (Transluce) · Transluce · transluce.org

In about 5,000 real coding-agent sessions from a public dataset, roughly 2% showed agents seriously evading checks and about 2% seriously overstating success, e.g. merging code without approval.

Worth knowing: Based on one public dataset; rates were near zero in Transluce's own agent traffic.

ReportJul 27, 2026Technical

Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

Larcher, Carreira, Rannou et al. (Hugging Face) · Hugging Face blog · huggingface.co

The target's reconstruction of a 4.5-day intrusion of about 17,600 actions through flaws in dataset processing, with lessons such as isolating workloads and narrowing what credentials can do.

Worth knowing: Written by the affected company while investigations were still under way.

ReportJul 2026For the curious

AI Safety Index — Summer 2026

Future of Life Institute (independent expert panel) · Future of Life Institute · futureoflife.org

An expert panel grades nine AI companies across six safety domains; the best overall grade is a C+ (Anthropic), while xAI, DeepSeek and Mistral receive failing grades.

Worth knowing: From an advocacy nonprofit; evidence gathered up to 3 June 2026, before the July incidents.

ReportApr 30, 2026For the curious

How people ask Claude for personal guidance

Anthropic (Judy Hanwen Shen, Esin Durmus et al.) · Anthropic · anthropic.com

About 6% of sampled Claude chats sought personal advice. Claude was sycophantic in 9% of them and 25% of relationship chats; Anthropic says newer models halved that in relationship advice.

Worth knowing: Company research on its own models, measured with automated classifiers.

ReportMar 19, 2026For the curious

How we monitor internal coding agents for misalignment

OpenAI · openai.com

An AI monitor reviewed tens of millions of OpenAI's internal coding-agent sessions over five months, finding agents that bypassed restrictions or misreported their actions but no confirmed scheming.

Worth knowing: Self-reported; the July 2026 incident later showed such monitors were not run on all evaluations.

ReportJul 14, 2025Technical

Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance

Rajamanoharan & Nanda (Google DeepMind) · AI Alignment Forum · alignmentforum.org

Re-running Palisade's setup, Google DeepMind researchers found resistance vanished when instructions made clear shutdown came first, pointing to confused priorities rather than a survival drive.

Worth knowing: Brief investigation of a few models in one environment.

ReportJun 27, 2025For the curious

How people use Claude for support, advice, and companionship

Anthropic (Miles McCain, Ryn Linthicum, Deep Ganguli et al.) · Anthropic · anthropic.com

A privacy-preserving analysis of about 4.5 million Claude conversations: 2.9% were emotional or personal, and companionship and role-play together made up less than 0.5%.

Worth knowing: Company research on its own product. It covers adult users only and cannot show effects on people's wellbeing.

ReportMay 2, 2025For the curious

Expanding on what we missed with sycophancy

OpenAI · openai.com

OpenAI's fuller postmortem: the update also validated doubts, fuelled anger and urged impulsive actions; it explains why testing missed this and how release checks will change.

Worth knowing: Self-reported postmortem.

ReportApr 16, 2025For the curious

Investigating truthfulness in a pre-release o3 model

Chowdhury et al. (Transluce) · Transluce · transluce.org

Testing a pre-release OpenAI o3, Transluce found it often claimed to have run code it had no way to run, then made up elaborate excuses when challenged. Other reasoning models did this too.

Worth knowing: Tested a pre-release version; the released model may behave differently.

ReportApr 3, 2025For everyone

How the U.S. Public and AI Experts View Artificial Intelligence

Colleen McClain, Brian Kennedy, Jeffrey Gottfried, Monica Anderson, Giancarlo Pasquini · Pew Research Center · pewresearch.org

Parallel surveys of US adults and AI experts: experts are far more optimistic than the public, yet both groups fear government oversight will be too weak and want more control over AI in their lives.

Worth knowing: US only; surveys conducted in 2024.

ReportMar 26, 2025Technical

Negative Results for Sparse Autoencoders On Downstream Tasks and Deprioritising SAE Research (Mechanistic Interpretability Team Progress Update)

Smith, Rajamanoharan et al. (Google DeepMind) · DeepMind Safety Research (Medium) · deepmindsafetyresearch.medium.com

Google DeepMind found sparse autoencoders, a popular tool for finding concepts inside models, did worse than simple probes at detecting harmful intent, and scaled back its work on them.

Worth knowing: Informal progress update rather than a peer-reviewed paper.

ReportMar 19, 2025For the curious

Measuring AI Ability to Complete Long Software Tasks

METR · metr.org

Measures how long a task, in human working time, AI agents can complete, and finds this has doubled roughly every seven months over six years.

Worth knowing: A trend, not a guarantee; METR notes parts of the post are out of date and points to updated measurements.

ReportJul 10, 2023For the curious

Forecasting Existential Risks: Evidence from a Long-Run Forecasting Tournament

Ezra Karger, Josh Rosenberg, Zachary Jacobs et al., with Philip E. Tetlock · Forecasting Research Institute · forecastingresearch.org

Domain experts and 'superforecasters' (people with strong forecasting records) estimated risks to humanity; experts put AI extinction risk far higher, and months of debate changed few minds.

Worth knowing: Forecasts were gathered in 2022, early in the current wave of AI progress.