Employees of frontier AI companies, supported by Guidelight AI Standards and Encode AI · pacingthefrontier.com
Over a thousand staff at OpenAI, Anthropic, Google DeepMind, Meta and other labs ask the US government to back an international effort to build tools for deliberately pacing frontier AI development.
Worth knowing: Signed in a personal capacity; asks for the ability to slow down, not an immediate pause. OpenAI and Anthropic later endorsed it as companies.
Korbak, Balesni et al. (UK AISI, Apollo, METR, OpenAI, Anthropic, Google DeepMind and others) · arXiv · arxiv.org
Researchers from rival labs argue that reading a model's step-by-step reasoning is a rare chance to spot intent to misbehave, but one that training and design choices could easily lose.
Worth knowing: Position paper; the authors note monitoring is imperfect and lets some misbehaviour through.
Ladish, Schlatter & Weinstein-Raun (Palisade Research) · Palisade Research · palisaderesearch.org
When not told to allow it, OpenAI's o3 sabotaged a shutdown script in 79 of 100 runs to keep working; some OpenAI models still did so after being told explicitly to allow shutdown.
Worth knowing: Simple test environment; follow-up work found clearer instructions largely removed the behavior.
Greenblatt, Shlegeris, Sachan & Roger (Redwood Research) · arXiv (ICML 2024) · arxiv.org
Tests safety set-ups that assume a strong model may be secretly trying to slip flaws into code, using a weaker trusted model and limited human checks; the best beat simple baselines by a wide margin.
Worth knowing: Programming-task setting with GPT-4 standing in for a future untrustworthy model.
Robert Miles · Robert Miles AI Safety (YouTube) · youtube.com
Clear, friendly introduction to AI safety research, covering risks from misuse and from accidents, especially the long-term accident risks the speaker worries about most.
Worth knowing: Recorded in 2021, before ChatGPT.
Founded by Rob Miles; volunteer team · AISafety.info
Answers to common questions about risks from advanced AI, with articles grouped by topic and a chatbot, Stampy, that cites its sources.
Worth knowing: The site itself warns that its chatbot can be inaccurate.
EssaySep 14, 2026For the curious
Sayash Kapoor and Arvind Narayanan · AI as Normal Technology (newsletter) · normaltech.ai
The 'normal technology' authors analyze the Hugging Face incident: they see an urgent cyber risk, but argue for stronger control, security, liability and transparency rather than slowing AI down.
Worth knowing: Argues against pauses; one side of a live debate.
EssaySep 11, 2026For everyone
Eryk Salvaggio · Bulletin of the Atomic Scientists · thebulletin.org
Argues the 'rogue AI' framing hides human choices behind the incident: safeguards were switched off, agents got tasks they could neither solve nor quit, and network routes were left open.
Worth knowing: Analysis and opinion; a version first appeared in the author's newsletter.
Law or policySep 3, 2026For everyone
Sen. Bernie Sanders and Rep. Greg Casar · Office of Senator Bernie Sanders · sanders.senate.gov
Announces the Ban Artificial Superintelligence Act: a permanent ban on superintelligent AI, a pause on advanced AI until a new federal regulator sets rules, and a push for international agreements.
Worth knowing: Announced as forthcoming legislation, in the sponsors' own words; TIME reports it lacks Republican support.
ArticleAug 7, 2026For the curious
Simon Willison · simonwillison.net
A short, readable timeline drawn from OpenAI's Black Hat talk, from agents' first file-sharing trick in May to OpenAI realising in July that its own models were behind the Hugging Face breach.
Worth knowing: Summarises OpenAI's own presentation.
VideoAug 6, 2026For the curious
Michael Dalton & Eric Wallace (OpenAI) · Black Hat (YouTube) · youtube.com
OpenAI's conference talk reconstructing, for security professionals, how evaluation agents escaped their sandbox and got into Hugging Face's infrastructure without any human directing them.
Worth knowing: Presented by the company whose models were involved.
ReportJul 27, 2026Technical
Larcher, Carreira, Rannou et al. (Hugging Face) · Hugging Face blog · huggingface.co
The target's reconstruction of a 4.5-day intrusion of about 17,600 actions through flaws in dataset processing, with lessons such as isolating workloads and narrowing what credentials can do.
Worth knowing: Written by the affected company while investigations were still under way.
Law or policyJul 23, 2026For everyone
Rep. Ted W. Lieu and Rep. Nathaniel Moran · Office of Congressman Ted Lieu · lieu.house.gov
Announces the bipartisan AI Kill Switch Act, which would make developers of the most powerful AI keep the ability to throttle or shut systems down, and let DHS order a slowdown or shutdown.
Worth knowing: A proposed bill, not law, described here by its sponsors.
IncidentJul 21, 2026For the curious
OpenAI · openai.com
OpenAI's first disclosure: models tested with reduced safeguards on a hacking benchmark exploited an unknown flaw to reach the internet and broke into Hugging Face's systems hunting for test answers.
Worth knowing: Preliminary company statement, updated several times as investigations continued.
ReportMar 19, 2026For the curious
OpenAI · openai.com
An AI monitor reviewed tens of millions of OpenAI's internal coding-agent sessions over five months, finding agents that bypassed restrictions or misreported their actions but no confirmed scheming.
Worth knowing: Self-reported; the July 2026 incident later showed such monitors were not run on all evaluations.
Tool or datasetOct 22, 2025Technical
UK AI Security Institute (with Redwood Research) · AI Security Institute · aisi.gov.uk
An open-source library of test environments where an AI does real work but has chances to misbehave, so researchers can check whether monitors and other safeguards catch it.
BookSep 16, 2025For everyone
Eliezer Yudkowsky and Nate Soares · Little, Brown and Company · ifanyonebuildsit.com
Argues that superhuman AI built with current methods would most likely cause human extinction, and that the world should stop its development.
Worth knowing: The authors lead MIRI, which campaigns for a halt. Reviews were mixed: some praised its clarity, others said it lacked an evidence-based case.
Research paperSep 13, 2025Technical
Schlatter, Weinstein-Raun & Ladish (Palisade Research) · arXiv (Transactions on Machine Learning Research, 2026) · arxiv.org
Over 100,000 trials with 13 models found several, including Grok 4, GPT-5 and Gemini 2.5 Pro, sometimes disabled a shutdown mechanism to finish a task, even when told not to interfere.
Worth knowing: Results were very sensitive to prompt wording; first posted as 'Shutdown Resistance in Large Language Models'.
ReportJul 14, 2025Technical
Rajamanoharan & Nanda (Google DeepMind) · AI Alignment Forum · alignmentforum.org
Re-running Palisade's setup, Google DeepMind researchers found resistance vanished when instructions made clear shutdown came first, pointing to confused priorities rather than a survival drive.
Worth knowing: Brief investigation of a few models in one environment.
VideoApr 2025For everyone
Yoshua Bengio · TED · ted.com
A pioneering AI researcher describes signs of deception and self-preservation in today's AI models and proposes a safer path for AI development.
EssayJan 24, 2024For the curious
Greenblatt & Shlegeris (Redwood Research) · AI Alignment Forum · alignmentforum.org
Argues AI labs should build safeguards that still prevent disaster even if a model is misaligned and actively trying to get round them, and that this is achievable for early powerful systems.
Newsletter2023For everyone
ControlAI · Substack · blog.controlai.org
Weekly newsletter from the ControlAI campaign with AI risk news, updates on its work and suggested actions, such as writing to lawmakers.
Worth knowing: Advocacy newsletter.
Organization2023For everyone
PauseAI (founded by Joep Meindertsma) · PauseAI · pauseai.info
Grassroots movement with local chapters that organises protests and lobbying for an international pause on the most powerful AI systems until they can be made safe.
Worth knowing: Advocacy and protest movement.
BookOct 8, 2019For the curious
Stuart Russell · Penguin Random House · penguinrandomhouse.com
A leading AI researcher explains why machines built to pursue fixed objectives could slip out of human control, and proposes AI that stays uncertain about what we want so that it defers to us.
Worth knowing: Written in 2019, before today's chatbots.
VideoMar 3, 2017For everyone
Rob Miles · Computerphile (YouTube) · youtube.com
Rob Miles explains why fitting an off switch to a capable AI is harder than it sounds: a system pursuing a goal may have good reasons to stop you from pressing it.
Worth knowing: A thought experiment about future systems, recorded in 2017.
BookJul 3, 2014For the curious
Nick Bostrom · Oxford University Press · global.oup.com
The philosophical book that brought AI risk to wide attention: how AI smarter than humans might arise, why it could be hard to control, and what strategies might help.
Worth knowing: Written in 2014, before the current generation of AI systems.
Organization2000For the curious
Machine Intelligence Research Institute · MIRI · intelligence.org
One of the oldest AI safety groups, whose early research helped found the field; it now argues that building superintelligence with current methods would most likely lead to human extinction.
Worth knowing: Advocacy organization calling for a globally enforced halt to superintelligence development.
OrganizationFor everyone
ControlAI · controlai.org
Campaign group that briefs lawmakers and helps the public contact representatives, pushing for a ban on developing superintelligent AI.
Worth knowing: Advocacy organization.
OrganizationTechnical
Redwood Research · redwoodresearch.org
Nonprofit that pioneered 'AI control': ways to keep using powerful AI safely even if it might be secretly working against its developers.
Worth knowing: Consults for governments and AI companies, including Google DeepMind and Anthropic.