Library

Everything worth reading about AI and human control. In one place.

Research papers, investigations, explainers, videos, laws and the organizations doing the work, from people who are alarmed and people who are skeptical. New additions are checked by two people before they are listed. The launch collection was compiled with the help of AI research assistants, and every link was opened and checked on 22 September 2026.

5 works · Tool or dataset · clear filters

Tool or datasetJul 2026For the curious

SaferAI Frontier Risk Management Tracker

SaferAI · tracker.safer-ai.org

Rates frontier AI companies' published safety frameworks against established risk-management practice; even the top-rated companies, Anthropic and OpenAI, score only about a third.

Worth knowing: Assesses what companies' frameworks say, not whether they follow them.

Tool or datasetMar 24, 2025For the curious

Introducing Docent

Meng, Huang, Steinhardt & Schwettmann (Transluce) · Transluce · transluce.org

A tool that uses AI to summarize, search and cluster long AI-agent transcripts, helping researchers spot broken tasks, unexpected behavior and weaknesses that a single score hides.

Tool or datasetApr 30, 2024For the curious

AI Lab Watch

Zach Stein-Perlman · AI Lab Watch · ailabwatch.org

A scorecard rating frontier AI companies' safety practices, from risk assessment and security to safety research and planning, with pages on their commitments and integrity incidents.

Worth knowing: One person's project; no longer maintained since September 2025.

Tool or datasetFor the curious

AISafety.com

AISafety.com

Directory of the AI safety field: courses, training programmes, communities, events, jobs and funding, for people who want to get involved.

Worth knowing: Framed around preventing human extinction from AI.

Tool or datasetFor the curious

Weval

The Collective Intelligence Project · Weval · weval.org

Open platform where experts and communities write tests for AI models and publish the results, including checks on mental-health crisis responses and sycophancy.

Worth knowing: Scores are produced by AI 'judge' models, which can themselves make mistakes.