Library

Everything worth reading about AI and human control. In one place.

Research papers, investigations, explainers, videos, laws and the organizations doing the work, from people who are alarmed and people who are skeptical. New additions are checked by two people before they are listed. The launch collection was compiled with the help of AI research assistants, and every link was opened and checked on 22 September 2026.

5 works on “Can we see what it’s thinking?” · clear filters

EssentialResearch paperJul 15, 2025Technical

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Korbak, Balesni et al. (UK AISI, Apollo, METR, OpenAI, Anthropic, Google DeepMind and others) · arXiv · arxiv.org

Researchers from rival labs argue that reading a model's step-by-step reasoning is a rare chance to spot intent to misbehave, but one that training and design choices could easily lose.

Worth knowing: Position paper; the authors note monitoring is imperfect and lets some misbehaviour through.

Research paperAug 20, 2026Technical

Scaling Activation Oracles to Trillion-Parameter Models

Choi et al. (Transluce) · Transluce · transluce.org

Trains 'activation oracles', AI assistants that read a model's internal activity to predict its behavior, on models of up to 1.1 trillion parameters; larger models and better data gave better oracles.

Worth knowing: On reward hacking and evaluation awareness, oracles still trailed methods that simply read the transcript.

ReportMar 26, 2025Technical

Negative Results for Sparse Autoencoders On Downstream Tasks and Deprioritising SAE Research (Mechanistic Interpretability Team Progress Update)

Smith, Rajamanoharan et al. (Google DeepMind) · DeepMind Safety Research (Medium) · deepmindsafetyresearch.medium.com

Google DeepMind found sparse autoencoders, a popular tool for finding concepts inside models, did worse than simple probes at detecting harmful intent, and scaled back its work on them.

Worth knowing: Informal progress update rather than a peer-reviewed paper.

Research paperMay 21, 2024Technical

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

Templeton et al. (Anthropic) · Transformer Circuits Thread · transformer-circuits.pub

Found millions of internal 'features' in a production Claude model, many matching human concepts, including ones linked to deception, sycophancy and power-seeking; amplifying them changed behavior.

Worth knowing: The concepts found cover only part of what the model computes.

Podcast2020Technical

AXRP - the AI X-risk Research Podcast

Daniel Filan · AXRP · axrp.net

Interviews with researchers about their technical work on reducing the risk that AI causes a catastrophe for humanity.

Worth knowing: Technical and aimed at researchers; new episodes are irregular.