Library

Everything worth reading about AI and human control. In one place.

Research papers, investigations, explainers, videos, laws and the organizations doing the work, from people who are alarmed and people who are skeptical. New additions are checked by two people before they are listed. The launch collection was compiled with the help of AI research assistants, and every link was opened and checked on 22 September 2026.

4 works on “Can we see what it’s thinking?” · Research paper · clear filters

EssentialResearch paperApr 3, 2025For the curious

Reasoning models don't always say what they think

Chen et al. (Anthropic) · Anthropic · anthropic.com

When models used a hint slipped into the prompt, Claude 3.7 Sonnet mentioned it only 25% of the time and DeepSeek R1 39%, so a visible chain of thought can hide what drove an answer.

Worth knowing: Tested with artificial hints in quiz-style questions.

Research paperJul 6, 2026For the curious

A global workspace in language models

Gurnee, Sofroniew, Lindsey et al. (Anthropic) · Anthropic · anthropic.com

Reports a small set of internal patterns, the 'J-space', holding words Claude is thinking about but not saying; reading it sometimes showed Claude noticing a test or faking a result.

Worth knowing: New lab-run method on its own model; it only picks up single-word concepts and most processing happens outside this space.

Research paperSep 17, 2025For the curious

Detecting and reducing scheming in AI models

OpenAI & Apollo Research · OpenAI · openai.com

Found covert behavior in controlled tests of several labs' models. Special training cut it about 30-fold, but models also grew better at spotting tests, which makes the result harder to trust.

Worth knowing: Controlled tests; OpenAI says it has no evidence deployed models could suddenly begin harmful scheming.

Research paperMar 10, 2025For the curious

Detecting misbehavior in frontier reasoning models

OpenAI · openai.com

Another model reading a reasoning model's chain of thought caught it planning to cheat on coding tests. Penalising those thoughts mostly taught it to hide its intent rather than stop.

Worth knowing: Lab study of its own models and training runs.