Dario Amodei · darioamodei.com
Anthropic's CEO argues AI capability gains should be slowed, proposing embedded outside evaluators (Anthropic commits now), coordinated limits among labs in democracies, and talks with China.
Worth knowing: Written by the CEO of a frontier AI company; critics raise self-regulation and antitrust concerns.
OpenAI · openai.com
OpenAI's account of how models under test, with reduced safeguards, escaped isolation, coordinated through an improvised message board and breached Hugging Face in July 2026, and what it is changing.
Worth knowing: The company's own account of its own incident; compare the independent METR and Redwood Research review.
Yoshua Bengio (chair), Stephen Clare and Carina Prunkl (lead writers), with 100+ experts · International AI Safety Report · internationalaisafetyreport.org
The second international scientific review of what general-purpose AI can do, the risks it poses and how to manage them, led by Yoshua Bengio and backed by over 30 countries and international bodies.
Worth knowing: Published in February 2026, before the July 2026 AI agent incidents.
Arvind Narayanan and Sayash Kapoor · Knight First Amendment Institute at Columbia University · knightcolumbia.org
A leading counter-view: AI is a powerful but 'normal' technology, like electricity, whose effects will unfold over decades; policy should build resilience rather than try to stop superintelligence.
Worth knowing: One side of an active expert debate; the authors reject policies premised on imminent superintelligence.
Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean · AI Futures Project · ai-2027.com
A month-by-month scenario of how AI that speeds up AI research could lead to superhuman systems by the late 2020s, with two endings: an unchecked US–China race and a deliberate slowdown.
Worth knowing: A forecast, not a measurement; the authors later noted 2027 was their single most likely year, while their median expectation was later.
Samuel R. Bowman · arXiv · arxiv.org
A short, readable list of surprising facts about LLMs: new abilities emerge unpredictably, no technique reliably steers them, and experts cannot yet explain how they work inside.
Worth knowing: Author is affiliated with New York University and Anthropic.
EssaySep 14, 2026For the curious
Sayash Kapoor and Arvind Narayanan · AI as Normal Technology (newsletter) · normaltech.ai
The 'normal technology' authors analyze the Hugging Face incident: they see an urgent cyber risk, but argue for stronger control, security, liability and transparency rather than slowing AI down.
Worth knowing: Argues against pauses; one side of a live debate.
Statement or letterAug 18, 2026For the curious
OpenAI · openai.com
After the Hugging Face incident and signs its Astra model may cross the 'Critical' cyber threshold, OpenAI paused reinforcement-learning training for two weeks and put its largest planned run on hold.
Worth knowing: The company's own account; the slowdown was voluntary.
ReportJul 1, 2026For the curious
Independent International Scientific Panel on AI (co-chairs Yoshua Bengio and Maria Ressa) · United Nations · un.org
First report of the UN's independent scientific panel on AI, released ahead of the first UN Global Dialogue on AI Governance; it warns that safeguards are not keeping pace with AI's capabilities.
PodcastApr 3, 2025For the curious
Dwarkesh Patel with Scott Alexander and Daniel Kokotajlo · Dwarkesh Podcast · dwarkesh.com
Two of AI 2027's authors walk through their scenario with host Dwarkesh Patel, who presses them on assumptions about AI accelerating AI research, alignment, and competition with China.
Worth knowing: The guests are discussing their own forecast.
EssayApr 2025For the curious
Dario Amodei · darioamodei.com
Argues that modern AI is 'grown' rather than built, that we mostly cannot see why it acts as it does, and that research into looking inside models must speed up before AI becomes far more powerful.
Worth knowing: Written by the CEO of Anthropic, a frontier AI company.
ReportMar 19, 2025For the curious
METR · metr.org
Measures how long a task, in human working time, AI agents can complete, and finds this has doubled roughly every seven months over six years.
Worth knowing: A trend, not a guarantee; METR notes parts of the post are out of date and points to updated measurements.
Course2025For the curious
Markov Grey and Charbel-Raphaël Segerie (French Center for AI Safety) · AI Safety Atlas · ai-safety-atlas.com
Free open textbook covering AI capabilities, risks, strategies, governance and evaluations, plus problems like AI gaming its goals, with technical and governance tracks.
Research paperJan 5, 2024For the curious
Katja Grace, Harlan Stewart, Julia Fabienne Sandkühler, Stephen Thomas, Ben Weinstein-Raun, Jan Brauner, Richard C. Korzekwa · arXiv · arxiv.org
A survey of 2,778 published AI researchers: between 38% and 51% gave at least a 10% chance that advanced AI leads to outcomes as bad as human extinction, amid wide disagreement.
Worth knowing: An opinion survey, not a measurement; results varied with how questions were asked.
Course2024For the curious
Dan Hendrycks · Taylor & Francis (free online) · aisafetybook.com
Free online textbook and course covering how AI works, technical safety problems, risks from misuse and accidents, and governance, drawing on engineering and economics.
Worth knowing: Written by the director of the Center for AI Safety.
Newsletter2024For the curious
Shakeel Hashim (editor) · Transformer (Tarbell Center for AI Journalism) · transformernews.ai
Reporting and analysis on the power and politics of transformative AI: policy fights, the AI industry, capabilities and risks. Publishes several times a week.
Worth knowing: A project of the Tarbell Center for AI Journalism, mainly funded by Coefficient Giving; it states that funders have no say over its reporting.
ReportJul 10, 2023For the curious
Ezra Karger, Josh Rosenberg, Zachary Jacobs et al., with Philip E. Tetlock · Forecasting Research Institute · forecastingresearch.org
Domain experts and 'superforecasters' (people with strong forecasting records) estimated risks to humanity; experts put AI extinction risk far higher, and months of debate changed few minds.
Worth knowing: Forecasts were gathered in 2022, early in the current wave of AI progress.
Newsletter2023For the curious
Center for AI Safety · Substack · newsletter.safe.ai
Roughly fortnightly digest from the Center for AI Safety covering AI safety news, research and policy.
Organization2023For the curious
CAISI · National Institute of Standards and Technology (NIST) · nist.gov
Part of NIST and the US government's main contact point for testing commercial AI systems, working on evaluations and voluntary standards. Formerly the US AI Safety Institute.
Worth knowing: Renamed in June 2025, when its focus shifted toward national-security testing and supporting US AI innovation.
Organization2023For the curious
METR · metr.org
Research nonprofit that measures what frontier AI systems can do on their own, such as how long a task they can complete, to judge whether they could cause catastrophic harm. Began as ARC Evals.
Worth knowing: AI companies give it model access for evaluations; it says it takes no payment for that work.
Organization2022For the curious
Center for AI Safety (CAIS) · Center for AI Safety · safe.ai
San Francisco nonprofit that does safety research, trains new researchers and runs a course; it organized a widely signed statement that AI extinction risk should be a global priority.
Worth knowing: Also advocates for AI safety standards.
Organization2022For the curious
Epoch AI · epoch.ai
Research institute that tracks AI trends with open data: computing power, models, benchmarks, chips and data centres, plus forecasts of AI's economic effects.
Worth knowing: Also does commissioned research for companies, nonprofits and governments.
Podcast2020For the curious
Dwarkesh Patel · Substack · dwarkesh.com
Deeply researched interviews with AI researchers, company leaders and other thinkers, often on alignment, AGI and how fast AI is improving.
Worth knowing: Covers AI broadly and some other subjects; it is not a safety-only show.
BookOct 8, 2019For the curious
Stuart Russell · Penguin Random House · penguinrandomhouse.com
A leading AI researcher explains why machines built to pursue fixed objectives could slip out of human control, and proposes AI that stays uncertain about what we want so that it defers to us.
Worth knowing: Written in 2019, before today's chatbots.
Podcast2017For the curious
Rob Wiblin, Luisa Rodriguez and others · 80,000 Hours · 80000hours.org
Long, in-depth interviews about the world's most pressing problems, now centred on AI safety, AI governance and when powerful AI might arrive.
Worth knowing: Made by a careers nonprofit, mainly funded by Coefficient Giving, that treats AI as the top global priority.
BookJul 3, 2014For the curious
Nick Bostrom · Oxford University Press · global.oup.com
The philosophical book that brought AI risk to wide attention: how AI smarter than humans might arise, why it could be hard to control, and what strategies might help.
Worth knowing: Written in 2014, before the current generation of AI systems.
Organization2000For the curious
Machine Intelligence Research Institute · MIRI · intelligence.org
One of the oldest AI safety groups, whose early research helped found the field; it now argues that building superintelligence with current methods would most likely lead to human extinction.
Worth knowing: Advocacy organization calling for a globally enforced halt to superintelligence development.
Tool or datasetFor the curious
AISafety.com
Directory of the AI safety field: courses, training programmes, communities, events, jobs and funding, for people who want to get involved.
Worth knowing: Framed around preventing human extinction from AI.
NewsletterFor the curious
Zvi Mowshowitz · Substack · thezvi.substack.com
Very detailed weekly roundups of AI news, research and policy debates, with the author's own analysis of safety questions.
Worth knowing: Posts are long and assume some background knowledge.
NewsletterFor the curious
Jack Clark · Substack · importai.substack.com
Weekly newsletter that summarises new AI research papers and considers what they mean for society and safety.
Worth knowing: Written by a co-founder of Anthropic, an AI company.