Responsible AI Collaborative · incidentdatabase.ai
Searchable collection of real-world cases where AI systems caused or nearly caused harm, modelled on incident records in aviation and computer security.
Worth knowing: Built from submitted reports, so it is not a complete count of AI harms.
Founded by Rob Miles; volunteer team · AISafety.info
Answers to common questions about risks from advanced AI, with articles grouped by topic and a chatbot, Stampy, that cites its sources.
Worth knowing: The site itself warns that its chatbot can be inaccurate.
Tool or datasetJul 2026For the curious
SaferAI · tracker.safer-ai.org
Rates frontier AI companies' published safety frameworks against established risk-management practice; even the top-rated companies, Anthropic and OpenAI, score only about a third.
Worth knowing: Assesses what companies' frameworks say, not whether they follow them.
Tool or dataset2026Technical
METR · metr.org
METR's index of the safety frameworks published by frontier AI companies, including Anthropic, OpenAI, Google DeepMind, Meta, xAI, Microsoft and Amazon, for comparing what each has committed to.
Worth knowing: The frameworks are written by the companies themselves.
Tool or datasetOct 22, 2025Technical
UK AI Security Institute (with Redwood Research) · AI Security Institute · aisi.gov.uk
An open-source library of test environments where an AI does real work but has chances to misbehave, so researchers can check whether monitors and other safeguards catch it.
Tool or datasetApr 2025For everyone
Damien Charlotin · damiencharlotin.com
A running record of court and tribunal decisions worldwide that found a party relied on AI-invented material, usually fake legal citations. It listed more than 2,000 cases by September 2026.
Worth knowing: Living database kept by one researcher; it counts only cases a court addressed, so the true number is higher.
Tool or datasetMar 24, 2025For the curious
Meng, Huang, Steinhardt & Schwettmann (Transluce) · Transluce · transluce.org
A tool that uses AI to summarize, search and cluster long AI-agent transcripts, helping researchers spot broken tasks, unexpected behavior and weaknesses that a single score hides.
Tool or datasetApr 30, 2024For the curious
Zach Stein-Perlman · AI Lab Watch · ailabwatch.org
A scorecard rating frontier AI companies' safety practices, from risk assessment and security to safety research and planning, with pages on their commitments and integrity incidents.
Worth knowing: One person's project; no longer maintained since September 2025.
Tool or datasetApr 2018For everyone
Victoria Krakovna and contributors · Google Sheets · docs.google.com
A crowd-sourced spreadsheet of real cases where AI systems found loopholes in the goals they were given, each with the intended goal, what the system did instead, and a source.
Worth knowing: Community-maintained list; many entries come from simple research or game settings.
Tool or datasetFor the curious
AISafety.com
Directory of the AI safety field: courses, training programmes, communities, events, jobs and funding, for people who want to get involved.
Worth knowing: Framed around preventing human extinction from AI.
Tool or datasetFor the curious
The Collective Intelligence Project · Weval · weval.org
Open platform where experts and communities write tests for AI models and publish the results, including checks on mental-health crisis responses and sycophancy.
Worth knowing: Scores are produced by AI 'judge' models, which can themselves make mistakes.