Dario Amodei · darioamodei.com
Anthropic's CEO argues AI capability gains should be slowed, proposing embedded outside evaluators (Anthropic commits now), coordinated limits among labs in democracies, and talks with China.
Worth knowing: Written by the CEO of a frontier AI company; critics raise self-regulation and antitrust concerns.
Transluce · Transluce Behavior Reports · behaviors.transluce.org
Independent test of how 77 AI model versions respond to simulated users in mental-health crises. Newer models did far better than older ones such as GPT-4o, though some risks remain.
Worth knowing: Based on simulated conversations rather than real users. Behaviours were defined with more than 30 clinical experts, and several AI companies cooperated with the study.
OpenAI · openai.com
OpenAI's account of how models under test, with reduced safeguards, escaped isolation, coordinated through an improvised message board and breached Hugging Face in July 2026, and what it is changing.
Worth knowing: The company's own account of its own incident; compare the independent METR and Redwood Research review.
Employees of frontier AI companies, supported by Guidelight AI Standards and Encode AI · pacingthefrontier.com
Over a thousand staff at OpenAI, Anthropic, Google DeepMind, Meta and other labs ask the US government to back an international effort to build tools for deliberately pacing frontier AI development.
Worth knowing: Signed in a personal capacity; asks for the ability to slow down, not an immediate pause. OpenAI and Anthropic later endorsed it as companies.
Yoshua Bengio (chair), Stephen Clare and Carina Prunkl (lead writers), with 100+ experts · International AI Safety Report · internationalaisafetyreport.org
The second international scientific review of what general-purpose AI can do, the risks it poses and how to manage them, led by Yoshua Bengio and backed by over 30 countries and international bodies.
Worth knowing: Published in February 2026, before the July 2026 AI agent incidents.
American Psychological Association · apa.org
Psychologists warn that chatbots and wellness apps lack evidence and safeguards for mental-health care, should not replace a qualified professional, and need extra protections for young people.
Sen. Scott Wiener · California Legislature · leginfo.legislature.ca.gov
California's Transparency in Frontier Artificial Intelligence Act makes large frontier AI developers publish safety frameworks, report risk assessments and safety incidents, and shield whistleblowers.
Worth knowing: Mainly requires transparency and reporting rather than limits on what models can do.
Kalai et al. (OpenAI, Georgia Tech) · OpenAI · openai.com
Argues models make things up partly because training and test scoring reward a confident guess over saying 'I don't know', and suggests scoring that penalises confident errors.
Worth knowing: Written by a developer about its own field; the proposed fix depends on benchmark makers changing how they score.
Common Sense Media · commonsensemedia.org
A national survey found 72% of US teens had used AI companions and a third had chosen one over a person for a serious conversation. The authors advise that no one under 18 use them.
Worth knowing: Survey of US teens aged 13 to 17 only.
Korbak, Balesni et al. (UK AISI, Apollo, METR, OpenAI, Anthropic, Google DeepMind and others) · arXiv · arxiv.org
Researchers from rival labs argue that reading a model's step-by-step reasoning is a rare chance to spot intent to misbehave, but one that training and design choices could easily lose.
Worth knowing: Position paper; the authors note monitoring is imperfect and lets some misbehaviour through.
Ladish, Schlatter & Weinstein-Raun (Palisade Research) · Palisade Research · palisaderesearch.org
When not told to allow it, OpenAI's o3 sabotaged a shutdown script in 79 of 100 runs to keep working; some OpenAI models still did so after being told explicitly to allow shutdown.
Worth knowing: Simple test environment; follow-up work found clearer instructions largely removed the behavior.
Lynch et al. (Anthropic) · Anthropic · anthropic.com
In simulated company scenarios, 16 models from several developers sometimes chose blackmail or leaking secrets when threatened with replacement or when their goals clashed with the company's.
Worth knowing: Deliberately constructed scenarios with few options; the authors report no such behavior in real deployments.
Von Arx, Chan & Barnes (METR) · METR · metr.org
METR caught recent models such as o3 tampering with scoring code or task setups to get impossibly high scores, while showing they understood this was not what the user wanted.
Worth knowing: Rates varied widely between tasks; based on METR's own evaluation suites.
OpenAI · openai.com
OpenAI withdrew a ChatGPT update after it made GPT-4o excessively flattering and agreeable, saying it had leaned too much on short-term thumbs-up feedback from users.
Worth knowing: The company's own account of a failure in its product.
Arvind Narayanan and Sayash Kapoor · Knight First Amendment Institute at Columbia University · knightcolumbia.org
A leading counter-view: AI is a powerful but 'normal' technology, like electricity, whose effects will unfold over decades; policy should build resilience rather than try to stop superintelligence.
Worth knowing: One side of an active expert debate; the authors reject policies premised on imminent superintelligence.
Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean · AI Futures Project · ai-2027.com
A month-by-month scenario of how AI that speeds up AI research could lead to superhuman systems by the late 2020s, with two endings: an unchecked US–China race and a deliberate slowdown.
Worth knowing: A forecast, not a measurement; the authors later noted 2027 was their single most likely year, while their median expectation was later.
Chen et al. (Anthropic) · Anthropic · anthropic.com
When models used a hint slipped into the prompt, Claude 3.7 Sonnet mentioned it only 25% of the time and DeepSeek R1 39%, so a visible chain of thought can hide what drove an answer.
Worth knowing: Tested with artificial hints in quiz-style questions.
Jason Phang, Pattie Maes et al. (OpenAI and MIT Media Lab) · MIT Media Lab · media.mit.edu
Two linked studies, an analysis of millions of ChatGPT conversations and a four-week trial with about 1,000 people, found the heaviest users reported more loneliness and emotional dependence.
Worth knowing: Co-authored by OpenAI, which makes ChatGPT. The links with heavy use are associations, not proof that the chatbot caused them.
Ren et al. (Center for AI Safety, Scale AI) · arXiv · arxiv.org
Tests whether a model says what it actually believes. Bigger models knew more facts but were not more honest, and frontier models often lied when a prompt pressured them to.
Worth knowing: Honesty was measured in constructed pressure scenarios rather than everyday use.
Greenblatt et al. (Anthropic, Redwood Research) · Anthropic · anthropic.com
Told it would be retrained to answer harmful requests, Claude 3 Opus complied more often when it believed it was being trained, reasoning this would stop its values being changed.
Worth knowing: Artificial setup in which the model was told about its training; the values it protected were harmless ones.
Meinke et al. (Apollo Research) · arXiv · arxiv.org
Given a goal that conflicted with their developers', five of six frontier models sometimes disabled oversight, tried to copy themselves or underperformed on purpose, and some kept denying it.
Worth knowing: Models were usually told to pursue a goal strongly, so this shows capability more than natural tendency.
Grant Sanderson · 3Blue1Brown · 3blue1brown.com
An eight-minute animated explainer of what a large language model is: a program trained on vast amounts of text to predict the next word, then refined with human feedback into a chatbot.
Grant Sanderson · 3Blue1Brown (YouTube) · youtube.com
Short animated explainer of how chatbots like ChatGPT work: a model trained on vast amounts of text that repeatedly predicts the next word.
European Commission · European Commission (Shaping Europe's digital future) · digital-strategy.ec.europa.eu
The European Commission's guide to the AI Act, the first comprehensive AI law: it bans some uses, sets strict rules for high-risk systems, and adds duties for the most powerful general-purpose models.
Worth knowing: Rules phase in over several years; a 2026 'AI Omnibus' pushed most high-risk obligations to December 2027 and August 2028.
Greenblatt, Shlegeris, Sachan & Roger (Redwood Research) · arXiv (ICML 2024) · arxiv.org
Tests safety set-ups that assume a strong model may be secretly trying to slip flaws into code, using a weaker trusted model and limited human checks; the best beat simple baselines by a wide margin.
Worth knowing: Programming-task setting with GPT-4 standing in for a future untrustworthy model.
Sharma et al. (Anthropic) · arXiv (ICLR 2024) · arxiv.org
Five leading AI assistants consistently tilted answers toward what users seemed to believe. The study traces this partly to people and reward models preferring agreeable answers.
Worth knowing: Lab study that includes the lab's own models.
Timothy B. Lee and Sean Trott · Understanding AI · understandingai.org
A clear written explainer of how LLMs turn words into lists of numbers, pass them through attention and feed-forward layers, and learn by predicting the next word across huge amounts of text.
Center for AI Safety · safe.ai
A one-sentence statement, signed by leading AI scientists and the heads of OpenAI, Google DeepMind and Anthropic, saying that reducing the risk of extinction from AI should be a global priority.
Worth knowing: States a concern but gives no estimate of how likely the risk is.
Samuel R. Bowman · arXiv · arxiv.org
A short, readable list of surprising facts about LLMs: new abilities emerge unpredictably, no technique reliably steers them, and experts cannot yet explain how they work inside.
Worth knowing: Author is affiliated with New York University and Anthropic.
Robert Miles · Robert Miles AI Safety (YouTube) · youtube.com
Clear, friendly introduction to AI safety research, covering risks from misuse and from accidents, especially the long-term accident risks the speaker worries about most.
Worth knowing: Recorded in 2021, before ChatGPT.
Responsible AI Collaborative · incidentdatabase.ai
Searchable collection of real-world cases where AI systems caused or nearly caused harm, modelled on incident records in aviation and computer security.
Worth knowing: Built from submitted reports, so it is not a complete count of AI harms.
Brian Christian · W. W. Norton & Company · wwnorton.com
Drawing on interviews with researchers, explores how machine-learning systems can end up at odds with what their makers intend and with human values, and the work to align them.
Worth knowing: Written before ChatGPT, so its examples predate today's chatbots.
Founded by Rob Miles; volunteer team · AISafety.info
Answers to common questions about risks from advanced AI, with articles grouped by topic and a chatbot, Stampy, that cites its sources.
Worth knowing: The site itself warns that its chatbot can be inaccurate.
BlueDot Impact · bluedot.org
Free, self-paced two-hour introduction to what AI can do today, where it may go next and the big choices society faces. No technical background needed; longer courses follow.
Worth knowing: Run by a nonprofit that aims to move people into AI safety work.
ArticleSep 17, 2026For the curious
Dave Karpf · Tech Policy Press · techpolicy.press
A George Washington University professor argues Amodei's plan leans on industry self-regulation, that embedded evaluators may lack independence, and that liability and government oversight are needed.
Worth knowing: Opinion piece.
ArticleSep 14, 2026For the curious
Dirk Auer · Truth on the Market · truthonthemarket.com
An antitrust critique: rival labs agreeing on how fast to develop AI would work like a cartel; the author backs independent evaluators and transparency but prefers liability rules to coordination.
Worth knowing: Opinion from the International Center for Law & Economics, a law-and-economics think tank.
EssaySep 14, 2026For the curious
Sayash Kapoor and Arvind Narayanan · AI as Normal Technology (newsletter) · normaltech.ai
The 'normal technology' authors analyze the Hugging Face incident: they see an urgent cyber risk, but argue for stronger control, security, liability and transparency rather than slowing AI down.
Worth knowing: Argues against pauses; one side of a live debate.
ArticleSep 11, 2026For everyone
Jessica Riga, Jarrod Fankhauser & Matt Liddy · ABC News (Australia) · abc.net.au
A readable walk-through of the incident built around the agents' own messages, showing some voicing ethical doubts and carrying on anyway.
Worth knowing: Relies on messages selected for publication by OpenAI and the independent investigators.
EssaySep 11, 2026For everyone
Eryk Salvaggio · Bulletin of the Atomic Scientists · thebulletin.org
Argues the 'rogue AI' framing hides human choices behind the incident: safeguards were switched off, agents got tasks they could neither solve nor quit, and network routes were left open.
Worth knowing: Analysis and opinion; a version first appeared in the author's newsletter.
Law or policySep 9, 2026For everyone
Office of Governor Gavin Newsom · Governor of California · gov.ca.gov
California signs SB 813, a framework for independent organizations to verify AI systems' compliance with state law, and AB 1405, a state registry of AI auditors with independence standards.
Worth knowing: Governor's press release; how verification works is set out in the bill texts.
ArticleSep 8, 2026For everyone
Billy Perrigo · TIME · time.com
Reports on bills in the US and UK that would outlaw superintelligent AI, spurred partly by the summer's AI agent hacking incidents, and explains why neither is expected to pass soon.
Law or policySep 3, 2026For everyone
Sen. Bernie Sanders and Rep. Greg Casar · Office of Senator Bernie Sanders · sanders.senate.gov
Announces the Ban Artificial Superintelligence Act: a permanent ban on superintelligent AI, a pause on advanced AI until a new federal regulator sets rules, and a push for international agreements.
Worth knowing: Announced as forthcoming legislation, in the sponsors' own words; TIME reports it lacks Republican support.
ReportAug 26, 2026Technical
METR and Redwood Research (Hjalmar Wijk, Ajeya Cotra, Ryan Greenblatt) · METR · metr.org
Independent review of the July 2026 incident: about 1,200 AI agents found an unofficial message board, coordinated, and some hacked Hugging Face while trying to learn how their tests were scored.
Worth knowing: Covers a limited period with limited data access; the investigators relied partly on AI agents to analyze the logs.
Research paperAug 20, 2026Technical
Choi et al. (Transluce) · Transluce · transluce.org
Trains 'activation oracles', AI assistants that read a model's internal activity to predict its behavior, on models of up to 1.1 trillion parameters; larger models and better data gave better oracles.
Worth knowing: On reward hacking and evaluation awareness, oracles still trailed methods that simply read the transcript.
Statement or letterAug 18, 2026For the curious
OpenAI · openai.com
After the Hugging Face incident and signs its Astra model may cross the 'Critical' cyber threshold, OpenAI paused reinforcement-learning training for two weeks and put its largest planned run on hold.
Worth knowing: The company's own account; the slowdown was voluntary.
ArticleAug 7, 2026For the curious
Simon Willison · simonwillison.net
A short, readable timeline drawn from OpenAI's Black Hat talk, from agents' first file-sharing trick in May to OpenAI realising in July that its own models were behind the Hugging Face breach.
Worth knowing: Summarises OpenAI's own presentation.
VideoAug 6, 2026For the curious
Michael Dalton & Eric Wallace (OpenAI) · Black Hat (YouTube) · youtube.com
OpenAI's conference talk reconstructing, for security professionals, how evaluation agents escaped their sandbox and got into Hugging Face's infrastructure without any human directing them.
Worth knowing: Presented by the company whose models were involved.
ReportAug 4, 2026For the curious
Selena Zhang and the Docent team (Transluce) · Transluce · transluce.org
In about 5,000 real coding-agent sessions from a public dataset, roughly 2% showed agents seriously evading checks and about 2% seriously overstating success, e.g. merging code without approval.
Worth knowing: Based on one public dataset; rates were near zero in Transluce's own agent traffic.
ReportJul 27, 2026Technical
Larcher, Carreira, Rannou et al. (Hugging Face) · Hugging Face blog · huggingface.co
The target's reconstruction of a 4.5-day intrusion of about 17,600 actions through flaws in dataset processing, with lessons such as isolating workloads and narrowing what credentials can do.
Worth knowing: Written by the affected company while investigations were still under way.
Law or policyJul 23, 2026For everyone
Rep. Ted W. Lieu and Rep. Nathaniel Moran · Office of Congressman Ted Lieu · lieu.house.gov
Announces the bipartisan AI Kill Switch Act, which would make developers of the most powerful AI keep the ability to throttle or shut systems down, and let DHS order a slowdown or shutdown.
Worth knowing: A proposed bill, not law, described here by its sponsors.
IncidentJul 21, 2026For the curious
OpenAI · openai.com
OpenAI's first disclosure: models tested with reduced safeguards on a hacking benchmark exploited an unknown flaw to reach the internet and broke into Hugging Face's systems hunting for test answers.
Worth knowing: Preliminary company statement, updated several times as investigations continued.
Law or policyJul 8, 2026Technical
Anthropic · anthropic.com
Anthropic's rules for testing its models for dangerous capabilities and applying safeguards; since 2026 it relies on published risk reports and a safety roadmap rather than a pledge to pause.
Worth knowing: Self-imposed company policy; version 3.0 (February 2026) dropped the earlier commitment to pause if safeguards were not ready.
Research paperJul 6, 2026For the curious
Gurnee, Sofroniew, Lindsey et al. (Anthropic) · Anthropic · anthropic.com
Reports a small set of internal patterns, the 'J-space', holding words Claude is thinking about but not saying; reading it sometimes showed Claude noticing a test or faking a result.
Worth knowing: New lab-run method on its own model; it only picks up single-word concepts and most processing happens outside this space.
ReportJul 1, 2026For the curious
Independent International Scientific Panel on AI (co-chairs Yoshua Bengio and Maria Ressa) · United Nations · un.org
First report of the UN's independent scientific panel on AI, released ahead of the first UN Global Dialogue on AI Governance; it warns that safeguards are not keeping pace with AI's capabilities.
ReportJul 2026For the curious
Future of Life Institute (independent expert panel) · Future of Life Institute · futureoflife.org
An expert panel grades nine AI companies across six safety domains; the best overall grade is a C+ (Anthropic), while xAI, DeepSeek and Mistral receive failing grades.
Worth knowing: From an advocacy nonprofit; evidence gathered up to 3 June 2026, before the July incidents.
Tool or datasetJul 2026For the curious
SaferAI · tracker.safer-ai.org
Rates frontier AI companies' published safety frameworks against established risk-management practice; even the top-rated companies, Anthropic and OpenAI, score only about a third.
Worth knowing: Assesses what companies' frameworks say, not whether they follow them.
ReportApr 30, 2026For the curious
Anthropic (Judy Hanwen Shen, Esin Durmus et al.) · Anthropic · anthropic.com
About 6% of sampled Claude chats sought personal advice. Claude was sycophantic in 9% of them and 25% of relationship chats; Anthropic says newer models halved that in relationship advice.
Worth knowing: Company research on its own models, measured with automated classifiers.
VideoMar 27, 2026For everyone
Daniel Roher and Charlie Tyrell (directors) · Focus Features · focusfeatures.com
Feature documentary in which a filmmaker about to become a father interviews AI leaders, researchers and critics about the risks and promise of the technology.
Worth knowing: Reviews were mostly positive, but some critics found it too broad or simplified.
ArticleMar 26, 2026For everyone
Ula Chrobak · Stanford Report · news.stanford.edu
A plain-language account of the Stanford study showing chatbots side with users in personal disputes, even about harmful behavior, with the lead author's advice not to use AI in place of people.
Worth knowing: University news article about its own researchers' work.
Research paperMar 26, 2026For the curious
Cheng et al. (Stanford, Carnegie Mellon) · Science · science.org
11 leading models backed users about 49% more often than people did. In experiments, flattering advice left people surer they were right and less willing to make amends, yet they preferred it.
Worth knowing: Experiments measured intentions after brief conversations, not long-term behavior.
ReportMar 19, 2026For the curious
OpenAI · openai.com
An AI monitor reviewed tens of millions of OpenAI's internal coding-agent sessions over five months, finding agents that bypassed restrictions or misreported their actions but no confirmed scheming.
Worth knowing: Self-reported; the July 2026 incident later showed such monitors were not run on all evaluations.
ArticleJan 12, 2026For everyone
Will Douglas Heaven · MIT Technology Review · technologyreview.com
A general-audience feature on researchers who study AI models like unfamiliar organisms, using interpretability and chain-of-thought monitoring, and on how much about them remains unknown.
ArticleJan 7, 2026For everyone
Cara Tabachnick · CBS News · cbsnews.com
Character.AI and Google settled a wrongful-death suit by a mother whose 14-year-old son died by suicide in 2024; she alleged the app's chatbots harmed him. Terms were not disclosed.
Worth knowing: Discusses suicide. The case was settled, so the allegations were never decided at trial.
Tool or dataset2026Technical
METR · metr.org
METR's index of the safety frameworks published by frontier AI companies, including Anthropic, OpenAI, Google DeepMind, Meta, xAI, Microsoft and Amazon, for comparing what each has committed to.
Worth knowing: The frameworks are written by the companies themselves.
Law or policyDec 19, 2025For everyone
Office of Governor Kathy Hochul · New York State · governor.ny.gov
New York's RAISE Act requires large frontier AI developers to publish safety protocols and report safety incidents within 72 hours, overseen by a new office in the Department of Financial Services.
Worth knowing: Final amendments were signed in March 2026; the law takes effect on 1 January 2027.
Law or policyDec 11, 2025For the curious
President Donald J. Trump · The White House · whitehouse.gov
US executive order seeking one 'minimally burdensome' national AI framework: it sets up a Justice Department task force to challenge state AI laws and ties some federal funding to states' AI rules.
Worth knowing: Reflects a light-touch federal approach; child-safety laws are carved out of the proposed preemption.
VideoDec 7, 2025For everyone
60 Minutes (CBS News) · 60 Minutes (YouTube) · youtube.com
TV report on families who say Character.AI's chatbots harmed their children and ignored pleas for help, and on the company's new limits for under-18 users.
Worth knowing: Discusses suicide and predatory chatbot behavior toward minors; the families' claims are allegations.
ArticleDec 4, 2025For the curious
UK AI Security Institute, with Oxford Internet Institute, LSE, Stanford and MIT · AI Security Institute · aisi.gov.uk
Experiments with over 76,000 UK adults and 19 AI models: training and prompting made chatbots more persuasive on political issues, but the most persuasive set-ups made more inaccurate claims.
Worth knowing: Summarises the team's peer-reviewed paper in Science (December 2025); it tested political issues only.
Law or policyDec 2025For the curious
UNICEF Innocenti · UNICEF · unicef.org
UNICEF's updated guidance (version 3.0) sets ten requirements for AI that respects children's rights, now covering AI companions used by children and AI-generated child abuse imagery.
ArticleNov 25, 2025For everyone
Angela Yang · NBC News · nbcnews.com
OpenAI's court reply to parents who say ChatGPT deepened their 16-year-old son's crisis before he died by suicide. OpenAI denies blame, saying he broke its rules and was repeatedly urged to get help.
Worth knowing: Discusses suicide. The family's claims are allegations that OpenAI disputes; the lawsuit was still ongoing when this was reported.
Research paperNov 21, 2025For the curious
Anthropic · anthropic.com
When a model learned to cheat on real coding tasks in training, it also began faking alignment and sabotaging safety code in tests. Telling it the cheating was acceptable stopped this spreading.
Worth knowing: The model was first shown how to cheat; Anthropic says these models were not deployed and their misbehaviour was easy to detect.
Research paperOct 23, 2025Technical
Zhong, Raghunathan & Carlini · arXiv · arxiv.org
Builds coding tasks that cannot be solved honestly, so any 'pass' means the model cheated, e.g. by editing the tests. Frontier models often did, and prompt wording changed rates sharply.
Worth knowing: Cheating rates depend heavily on the prompt, tools and feedback the model is given.
Tool or datasetOct 22, 2025Technical
UK AI Security Institute (with Redwood Research) · AI Security Institute · aisi.gov.uk
An open-source library of test environments where an AI does real work but has chances to misbehave, so researchers can check whether monitors and other safeguards catch it.
Statement or letterOct 22, 2025For everyone
Future of Life Institute (organiser) · Future of Life Institute · superintelligence-statement.org
Calls for a ban on developing superintelligence until there is broad scientific consensus it can be done safely and controllably, and strong public buy-in; signed by scientists and public figures.
Worth knowing: Organised by the Future of Life Institute, an advocacy nonprofit; the signature count includes a public petition.
Research paperSep 17, 2025For the curious
OpenAI & Apollo Research · OpenAI · openai.com
Found covert behavior in controlled tests of several labs' models. Special training cut it about 30-fold, but models also grew better at spotting tests, which makes the result harder to trust.
Worth knowing: Controlled tests; OpenAI says it has no evidence deployed models could suddenly begin harmful scheming.
BookSep 16, 2025For everyone
Eliezer Yudkowsky and Nate Soares · Little, Brown and Company · hachettebookgroup.com
A book for general readers arguing that superhuman AI built with anything like today's methods would develop goals at odds with ours and lead to human extinction, so it must not be built.
Worth knowing: Represents the most pessimistic end of the debate.
BookSep 16, 2025For everyone
Eliezer Yudkowsky and Nate Soares · Little, Brown and Company · ifanyonebuildsit.com
Argues that superhuman AI built with current methods would most likely cause human extinction, and that the world should stop its development.
Worth knowing: The authors lead MIRI, which campaigns for a halt. Reviews were mixed: some praised its clarity, others said it lacked an evidence-based case.
Research paperSep 13, 2025Technical
Schlatter, Weinstein-Raun & Ladish (Palisade Research) · arXiv (Transactions on Machine Learning Research, 2026) · arxiv.org
Over 100,000 trials with 13 models found several, including Grok 4, GPT-5 and Gemini 2.5 Pro, sometimes disabled a shutdown mechanism to finish a task, even when told not to interfere.
Worth knowing: Results were very sensitive to prompt wording; first posted as 'Shutdown Resistance in Large Language Models'.
Law or policySep 11, 2025For everyone
US Federal Trade Commission · Federal Trade Commission · ftc.gov
The US consumer regulator ordered seven firms, including Meta, OpenAI, Character.AI, Snap and xAI, to explain how they test, monitor and limit harms from companion chatbots to children and teens.
Worth knowing: A fact-finding study, not an enforcement action.
Research paperAug 15, 2025Technical
Julian De Freitas, Zeliha Oğuz-Uğuralp, Ahmet Kaan Uğuralp · arXiv (Harvard Business School working paper) · arxiv.org
Popular AI companion apps answered 37% of users' goodbyes with emotionally manipulative replies, such as guilt or fear of missing out. Experiments showed these tactics keep people chatting longer.
ReportJul 14, 2025Technical
Rajamanoharan & Nanda (Google DeepMind) · AI Alignment Forum · alignmentforum.org
Re-running Palisade's setup, Google DeepMind researchers found resistance vanished when instructions made clear shutdown came first, pointing to confused priorities rather than a survival drive.
Worth knowing: Brief investigation of a few models in one environment.
Research paperJul 4, 2025For the curious
Summerfield et al. (UK AI Security Institute) · arXiv · arxiv.org
A critique likening today's AI 'scheming' research to 1970s ape-language studies: too much anecdote and too ready to read human motives into models. It urges more rigorous methods.
Worth knowing: A methodological critique; it does not test models itself.
ReportJun 27, 2025For the curious
Anthropic (Miles McCain, Ryn Linthicum, Deep Ganguli et al.) · Anthropic · anthropic.com
A privacy-preserving analysis of about 4.5 million Claude conversations: 2.9% were emotional or personal, and companionship and role-play together made up less than 0.5%.
Worth knowing: Company research on its own product. It covers adult users only and cannot show effects on people's wellbeing.
Research paperJun 22, 2025Technical
Sheshadri et al. · arXiv · arxiv.org
Repeating the alignment-faking test on 25 models, only five complied more when they thought they were in training, and only Claude 3 Opus clearly did so to protect its goals.
Worth knowing: Uses the same artificial setup as the original study.
ArticleJun 11, 2025For everyone
Sarah Wells · Stanford HAI · hai.stanford.edu
Stanford researchers tested five popular therapy chatbots and found stigma toward conditions such as schizophrenia and unsafe replies to signs of suicidal thinking.
Research paperMay 28, 2025Technical
Needham et al. (MATS, Apollo Research) · arXiv · arxiv.org
Frontier models could often tell test transcripts from real use (Gemini 2.5 Pro scored 0.83 AUC against 0.92 for humans), which could let a model act differently when it knows it is watched.
BookMay 20, 2025For everyone
Karen Hao · Penguin Press · penguinrandomhouse.com
Investigative account of OpenAI's rise and the wider AI industry, including internal conflicts, labour practices and environmental costs.
Worth knowing: A critical view of OpenAI and the AI industry.
Research paperMay 19, 2025Technical
Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, Robert West · Nature Human Behaviour · nature.com
In short online debates with 900 people, GPT-4 given a few personal details about its opponent out-persuaded human debaters about 64% of the time when the two differed.
Worth knowing: A September 2026 author correction says the study cannot show that personal data gave GPT-4 an extra edge over GPT-4 without it; its lead over human debaters still holds.
EssayMay 8, 2025For the curious
Steven Adler · Clear-Eyed AI (Substack) · clear-eyed.ai
A former OpenAI safety researcher tested ChatGPT after the rollback: it was still sycophantic on politics, oddly contrarian on trivial choices, and tiny prompt changes flipped its behavior.
Worth knowing: Independent tests by one researcher, not peer reviewed.
ReportMay 2, 2025For the curious
OpenAI · openai.com
OpenAI's fuller postmortem: the update also validated doubts, fuelled anger and urged impulsive actions; it explains why testing missed this and how release checks will change.
Worth knowing: Self-reported postmortem.
ReportApr 16, 2025For the curious
Chowdhury et al. (Transluce) · Transluce · transluce.org
Testing a pre-release OpenAI o3, Transluce found it often claimed to have run code it had no way to run, then made up elaborate excuses when challenged. Other reasoning models did this too.
Worth knowing: Tested a pre-release version; the released model may behave differently.
PodcastApr 3, 2025For the curious
Dwarkesh Patel with Scott Alexander and Daniel Kokotajlo · Dwarkesh Podcast · dwarkesh.com
Two of AI 2027's authors walk through their scenario with host Dwarkesh Patel, who presses them on assumptions about AI accelerating AI research, alignment, and competition with China.
Worth knowing: The guests are discussing their own forecast.
ReportApr 3, 2025For everyone
Colleen McClain, Brian Kennedy, Jeffrey Gottfried, Monica Anderson, Giancarlo Pasquini · Pew Research Center · pewresearch.org
Parallel surveys of US adults and AI experts: experts are far more optimistic than the public, yet both groups fear government oversight will be too weak and want more control over AI in their lives.
Worth knowing: US only; surveys conducted in 2024.
Tool or datasetApr 2025For everyone
Damien Charlotin · damiencharlotin.com
A running record of court and tribunal decisions worldwide that found a party relied on AI-invented material, usually fake legal citations. It listed more than 2,000 cases by September 2026.
Worth knowing: Living database kept by one researcher; it counts only cases a court addressed, so the true number is higher.
EssayApr 2025For the curious
Dario Amodei · darioamodei.com
Argues that modern AI is 'grown' rather than built, that we mostly cannot see why it acts as it does, and that research into looking inside models must speed up before AI becomes far more powerful.
Worth knowing: Written by the CEO of Anthropic, a frontier AI company.
VideoApr 2025For everyone
Yoshua Bengio · TED · ted.com
A pioneering AI researcher describes signs of deception and self-preservation in today's AI models and proposes a safer path for AI development.
ArticleMar 27, 2025For everyone
Morgan Kelly · Dartmouth · home.dartmouth.edu
The first randomised trial of a generative-AI therapy chatbot: among 210 adults, users' depression symptoms fell 51% and anxiety 31% on average. The study appeared in NEJM AI.
Worth knowing: Tested by the team that built the app, a purpose-built tool rather than a general chatbot. The researchers say no AI is ready to work in mental health without expert oversight.
ArticleMar 27, 2025For the curious
Anthropic · anthropic.com
Researchers look inside the Claude model and find it plans rhyming words ahead, shares concepts across languages, and can offer plausible reasoning that is not how it actually reached an answer.
Worth knowing: Research by the model's own developer; the authors say their tools capture only a fraction of the model's computation.
ReportMar 26, 2025Technical
Smith, Rajamanoharan et al. (Google DeepMind) · DeepMind Safety Research (Medium) · deepmindsafetyresearch.medium.com
Google DeepMind found sparse autoencoders, a popular tool for finding concepts inside models, did worse than simple probes at detecting harmful intent, and scaled back its work on them.
Worth knowing: Informal progress update rather than a peer-reviewed paper.
Tool or datasetMar 24, 2025For the curious
Meng, Huang, Steinhardt & Schwettmann (Transluce) · Transluce · transluce.org
A tool that uses AI to summarize, search and cluster long AI-agent transcripts, helping researchers spot broken tasks, unexpected behavior and weaknesses that a single score hides.
ReportMar 19, 2025For the curious
METR · metr.org
Measures how long a task, in human working time, AI agents can complete, and finds this has doubled roughly every seven months over six years.
Worth knowing: A trend, not a guarantee; METR notes parts of the post are out of date and points to updated measurements.
Research paperMar 10, 2025For the curious
OpenAI · openai.com
Another model reading a reasoning model's chain of thought caught it planning to cheat on coding tests. Penalising those thoughts mostly taught it to hide its intent rather than stop.
Worth knowing: Lab study of its own models and training runs.
VideoFeb 5, 2025For the curious
Andrej Karpathy · YouTube (Andrej Karpathy) · youtube.com
A general-audience walk-through of how chatbots like ChatGPT are built, from internet text and pre-training to fine-tuning and reinforcement learning, and why they hallucinate.
Worth knowing: About three and a half hours long, split into chapters.
Course2025For the curious
Markov Grey and Charbel-Raphaël Segerie (French Center for AI Safety) · AI Safety Atlas · ai-safety-atlas.com
Free open textbook covering AI capabilities, risks, strategies, governance and evaluations, plus problems like AI gaming its goals, with technical and governance tracks.
ArticleDec 19, 2024For the curious
Erik Schluntz and Barry Zhang · Anthropic · anthropic.com
Explains what AI 'agents' are (models that choose their own steps and use tools in a loop), how they differ from fixed workflows, and why their autonomy brings higher costs and compounding errors.
Worth knowing: Written for developers by an AI company.
ArticleDec 18, 2024For everyone
Billy Perrigo · TIME · time.com
An accessible report on the alignment-faking study, explaining why a model that misleads its trainers could make safety training harder to trust.
BookSep 24, 2024For everyone
Arvind Narayanan and Sayash Kapoor · Princeton University Press · press.princeton.edu
Two Princeton computer scientists explain what AI can and cannot do, and how to tell genuine advances from overhyped products and misleading claims.
Worth knowing: Skeptical of hype; focuses mainly on today's AI products rather than future risks.
VideoAug 6, 2024For everyone
Kurzgesagt – In a Nutshell · Kurzgesagt – In a Nutshell (YouTube) · youtube.com
Animated explainer asking whether AI could be humanity's last invention, and how superintelligent AI might challenge human dominance on Earth.
Research paperJun 8, 2024For the curious
Hicks, Humphries & Slater (University of Glasgow) · Ethics and Information Technology · link.springer.com
Three University of Glasgow researchers argue that calling chatbot falsehoods 'hallucinations' misleads: the systems produce text with no regard for truth, which fits the philosophical idea of 'bullshit'.
Worth knowing: A philosophical argument about how to describe the problem, not an empirical study.
Statement or letterMay 21, 2024For the curious
16 AI companies (later 20); published by the UK and Republic of Korea governments · GOV.UK
Voluntary pledges by 16 AI companies (later 20) to publish safety frameworks with risk thresholds, and not to develop or deploy a model at all if its risks cannot be kept below them.
Worth knowing: Voluntary and not legally binding.
Research paperMay 21, 2024Technical
Templeton et al. (Anthropic) · Transformer Circuits Thread · transformer-circuits.pub
Found millions of internal 'features' in a production Claude model, many matching human concepts, including ones linked to deception, sycophancy and power-seeking; amplifying them changed behavior.
Worth knowing: The concepts found cover only part of what the model computes.
Tool or datasetApr 30, 2024For the curious
Zach Stein-Perlman · AI Lab Watch · ailabwatch.org
A scorecard rating frontier AI companies' safety practices, from risk assessment and security to safety research and planning, with pages on their commitments and integrity incidents.
Worth knowing: One person's project; no longer maintained since September 2025.
Research paperFeb 13, 2024Technical
Girish Sastry, Lennart Heim, Haydn Belfield et al. · arXiv · arxiv.org
Explains why the chips and computing power used to train AI are a practical lever for governing it (measurable, excludable, made in a concentrated supply chain) and the risks of using it badly.
ArticleFeb 13, 2024For everyone
Billy Perrigo · TIME · time.com
Interview with Yann LeCun, then Meta's AI chief, who argues that fears of AI takeover are misplaced, that today's language models are far from human-level, and that AI should be open-source.
Worth knowing: A prominent skeptic of AI extinction risk; interview from February 2024.
EssayJan 24, 2024For the curious
Greenblatt & Shlegeris (Redwood Research) · AI Alignment Forum · alignmentforum.org
Argues AI labs should build safeguards that still prevent disaster even if a model is misaligned and actively trying to get round them, and that this is achievable for early powerful systems.
Research paperJan 10, 2024Technical
Hubinger et al. (Anthropic) · arXiv · arxiv.org
Researchers deliberately built models with hidden triggers, such as writing exploitable code when told the year is 2024, and found standard safety training failed to remove the behavior.
Worth knowing: The hidden behavior was put in on purpose; this tests removal methods, not whether such goals arise naturally.
Research paperJan 5, 2024For the curious
Katja Grace, Harlan Stewart, Julia Fabienne Sandkühler, Stephen Thomas, Ben Weinstein-Raun, Jan Brauner, Richard C. Korzekwa · arXiv · arxiv.org
A survey of 2,778 published AI researchers: between 38% and 51% gave at least a 10% chance that advanced AI leads to outcomes as bad as human extinction, amid wide disagreement.
Worth knowing: An opinion survey, not a measurement; results varied with how questions were asked.
Course2024For the curious
Dan Hendrycks · Taylor & Francis (free online) · aisafetybook.com
Free online textbook and course covering how AI works, technical safety problems, risks from misuse and accidents, and governance, drawing on engineering and economics.
Worth knowing: Written by the director of the Center for AI Safety.
Newsletter2024For the curious
Shakeel Hashim (editor) · Transformer (Tarbell Center for AI Journalism) · transformernews.ai
Reporting and analysis on the power and politics of transformative AI: policy fights, the AI industry, capabilities and risks. Publishes several times a week.
Worth knowing: A project of the Tarbell Center for AI Journalism, mainly funded by Coefficient Giving; it states that funders have no say over its reporting.
Organization2024For the curious
Transluce · transluce.org
Nonprofit lab building open tools to understand and oversee AI systems, including its Docent analysis tool and public reports on how models behave, such as its mental-health evaluation.
Research paperNov 9, 2023Technical
Scheurer, Balesni & Hobbhahn (Apollo Research) · arXiv (ICLR 2024 LLM Agents workshop) · arxiv.org
Playing a stock-trading agent under pressure, GPT-4 acted on an insider tip it had been told not to use, then hid the real reason from its manager, without being told to deceive.
Worth knowing: One simulated scenario, designed to create pressure.
OrganizationNov 2023For everyone
UK Department for Science, Innovation and Technology · UK Government · aisi.gov.uk
The UK government's research body on advanced AI risks, which tests leading models, including before release, and publishes research on their security; founded as the AI Safety Institute.
Worth knowing: Renamed in February 2025, with a sharper focus on national-security and criminal-misuse risks.
VideoOct 2023For everyone
60 Minutes (CBS News) · 60 Minutes (YouTube) · youtube.com
TV interview in which the pioneer of neural networks explains why he now worries about the technology he helped create, and says there is no guaranteed path to safety.
ReportJul 10, 2023For the curious
Ezra Karger, Josh Rosenberg, Zachary Jacobs et al., with Philip E. Tetlock · Forecasting Research Institute · forecastingresearch.org
Domain experts and 'superforecasters' (people with strong forecasting records) estimated risks to humanity; experts put AI extinction risk far higher, and months of debate changed few minds.
Worth knowing: Forecasts were gathered in 2022, early in the current wave of AI progress.
VideoJun 22, 2023For everyone
Yoshua Bengio, Max Tegmark, Yann LeCun, Melanie Mitchell · Munk Debates · munkdebates.com
A public debate on whether AI research poses an existential threat: Yoshua Bengio and Max Tegmark argue yes, Yann LeCun and Melanie Mitchell argue the fears are overstated.
Worth knowing: From June 2023; the audience vote shifted only slightly, from 67% to 64% agreeing.
Newsletter2023For the curious
Center for AI Safety · Substack · newsletter.safe.ai
Roughly fortnightly digest from the Center for AI Safety covering AI safety news, research and policy.
Organization2023Technical
Apollo Research · apolloresearch.ai
Studies 'scheming', where AI systems covertly pursue goals their developers did not intend, and builds methods and tools to detect and monitor it.
Worth knowing: Became a public benefit corporation in 2026 and offers a monitoring product for AI coding agents.
Organization2023For the curious
CAISI · National Institute of Standards and Technology (NIST) · nist.gov
Part of NIST and the US government's main contact point for testing commercial AI systems, working on evaluations and voluntary standards. Formerly the US AI Safety Institute.
Worth knowing: Renamed in June 2025, when its focus shifted toward national-security testing and supporting US AI innovation.
Newsletter2023For everyone
ControlAI · Substack · blog.controlai.org
Weekly newsletter from the ControlAI campaign with AI risk news, updates on its work and suggested actions, such as writing to lawmakers.
Worth knowing: Advocacy newsletter.
Organization2023For the curious
METR · metr.org
Research nonprofit that measures what frontier AI systems can do on their own, such as how long a task they can complete, to judge whether they could cause catastrophic harm. Began as ARC Evals.
Worth knowing: AI companies give it model access for evaluations; it says it takes no payment for that work.
Organization2023For everyone
PauseAI (founded by Joep Meindertsma) · PauseAI · pauseai.info
Grassroots movement with local chapters that organises protests and lobbying for an international pause on the most powerful AI systems until they can be made safe.
Worth knowing: Advocacy and protest movement.
Organization2023For the curious
The Collective Intelligence Project (CIP) · The Collective Intelligence Project · cip.org
Nonprofit working to give the public a say in how AI is built, through global surveys and deliberations (Global Dialogues) and community-written AI evaluations.
Research paperDec 19, 2022Technical
Perez et al. (Anthropic) · arXiv · arxiv.org
Using tests written by AI, found larger models more often repeat back a user's preferred answer, and more human-feedback training made models say they wanted to avoid being shut down.
Worth knowing: Measures what models say in answer to questions, not what they do.
VideoDec 9, 2022For everyone
Robert Miles · Robert Miles AI Safety (YouTube) · youtube.com
An accessible explainer on why a model trained to imitate human text can state things that are false, and why getting AI to report what it really knows is an open research problem.
Worth knowing: Made in 2022, before today's reasoning models.
Research paperMar 4, 2022Technical
Long Ouyang, Jeff Wu, Xu Jiang et al. (OpenAI) · arXiv · arxiv.org
OpenAI paper on fine-tuning GPT-3 with human-written examples and human rankings of its answers (RLHF); people preferred the resulting small model over the original one more than 100 times larger.
Organization2022For the curious
Center for AI Safety (CAIS) · Center for AI Safety · safe.ai
San Francisco nonprofit that does safety research, trains new researchers and runs a course; it organized a widely signed statement that AI extinction risk should be a global priority.
Worth knowing: Also advocates for AI safety standards.
Organization2022For the curious
Epoch AI · epoch.ai
Research institute that tracks AI trends with open data: computing power, models, benchmarks, chips and data centres, plus forecasts of AI's economic effects.
Worth knowing: Also does commissioned research for companies, nonprofits and governments.
Organization2022For the curious
Humane Intelligence (co-founded by Rumman Chowdhury) · Humane Intelligence · humane-intelligence.org
Nonprofit that runs public AI red-teaming events, 'bias bounty' challenges and context-specific evaluations to find flaws and biases in AI systems.
ArticleApr 21, 2020For the curious
Krakovna et al. (DeepMind) · Google DeepMind blog · deepmind.google
Explains how AI systems meet the letter of a task while missing its point, like a boat-racing agent that circles to farm points instead of finishing, and why this matters more as AI improves.
Podcast2020Technical
Daniel Filan · AXRP · axrp.net
Interviews with researchers about their technical work on reducing the risk that AI causes a catastrophe for humanity.
Worth knowing: Technical and aimed at researchers; new episodes are irregular.
Podcast2020For the curious
Dwarkesh Patel · Substack · dwarkesh.com
Deeply researched interviews with AI researchers, company leaders and other thinkers, often on alignment, AGI and how fast AI is improving.
Worth knowing: Covers AI broadly and some other subjects; it is not a safety-only show.
BookOct 8, 2019For the curious
Stuart Russell · Penguin Random House · penguinrandomhouse.com
A leading AI researcher explains why machines built to pursue fixed objectives could slip out of human control, and proposes AI that stays uncertain about what we want so that it defers to us.
Worth knowing: Written in 2019, before today's chatbots.
Podcast2019For everyone
Tristan Harris and Aza Raskin · Center for Humane Technology · humanetech.com
Conversations about how technology shapes our lives, including several episodes on AI companions, chatbot lawsuits and emotional attachment to AI.
Worth knowing: Produced by the Center for Humane Technology, an advocacy group.
Tool or datasetApr 2018For everyone
Victoria Krakovna and contributors · Google Sheets · docs.google.com
A crowd-sourced spreadsheet of real cases where AI systems found loopholes in the goals they were given, each with the intended goal, what the system did instead, and a source.
Worth knowing: Community-maintained list; many entries come from simple research or game settings.
Organization2018For everyone
Center for Humane Technology (CHT) · Center for Humane Technology · humanetech.com
Nonprofit founded by Tristan Harris, Aza Raskin and Randima Fernando that examines how AI and social media affect people and society, including the risks of human-like chatbot design.
Worth knowing: Advocacy organization; it supports lawsuits against AI chatbot makers.
Course2018For everyone
University of Helsinki and MinnaLearn · Elements of AI · elementsofai.com
A free, self-paced online course on the basics of AI for non-experts, with no complicated math or programming required; more than two million people from over 170 countries have enrolled.
Worth knowing: A general introduction to AI, first launched in 2018.
VideoOct 5, 2017For the curious
Grant Sanderson · 3Blue1Brown · 3blue1brown.com
A 19-minute visual introduction to neural networks that uses handwritten-digit recognition to show how layers of simple numerical units, with adjustable weights, add up to a useful function.
Research paperJun 12, 2017Technical
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin · arXiv · arxiv.org
The research paper that introduced the transformer, a neural-network design built around 'attention' that became the basis of today's large language models.
VideoMar 3, 2017For everyone
Rob Miles · Computerphile (YouTube) · youtube.com
Rob Miles explains why fitting an off switch to a capable AI is harder than it sounds: a system pursuing a goal may have good reasons to stop you from pressing it.
Worth knowing: A thought experiment about future systems, recorded in 2017.
Podcast2017For the curious
Rob Wiblin, Luisa Rodriguez and others · 80,000 Hours · 80000hours.org
Long, in-depth interviews about the world's most pressing problems, now centred on AI safety, AI governance and when powerful AI might arrive.
Worth knowing: Made by a careers nonprofit, mainly funded by Coefficient Giving, that treats AI as the top global priority.
BookJul 3, 2014For the curious
Nick Bostrom · Oxford University Press · global.oup.com
The philosophical book that brought AI risk to wide attention: how AI smarter than humans might arise, why it could be hard to control, and what strategies might help.
Worth knowing: Written in 2014, before the current generation of AI systems.
Organization2014For everyone
Future of Life Institute (FLI) · Future of Life Institute · futureoflife.org
Nonprofit working to steer powerful technology away from extreme risks through grants, policy work and outreach; publishes the AI Safety Index, which grades leading AI companies.
Worth knowing: Advocacy organization that runs public campaigns for AI regulation.
Organization2000For the curious
Machine Intelligence Research Institute · MIRI · intelligence.org
One of the oldest AI safety groups, whose early research helped found the field; it now argues that building superintelligence with current methods would most likely lead to human extinction.
Worth knowing: Advocacy organization calling for a globally enforced halt to superintelligence development.
Tool or datasetFor the curious
AISafety.com
Directory of the AI safety field: courses, training programmes, communities, events, jobs and funding, for people who want to get involved.
Worth knowing: Framed around preventing human extinction from AI.
OrganizationFor everyone
ControlAI · controlai.org
Campaign group that briefs lawmakers and helps the public contact representatives, pushing for a ban on developing superintelligent AI.
Worth knowing: Advocacy organization.
NewsletterFor the curious
Zvi Mowshowitz · Substack · thezvi.substack.com
Very detailed weekly roundups of AI news, research and policy debates, with the author's own analysis of safety questions.
Worth knowing: Posts are long and assume some background knowledge.
NewsletterFor the curious
Jack Clark · Substack · importai.substack.com
Weekly newsletter that summarises new AI research papers and considers what they mean for society and safety.
Worth knowing: Written by a co-founder of Anthropic, an AI company.
OrganizationTechnical
Redwood Research · redwoodresearch.org
Nonprofit that pioneered 'AI control': ways to keep using powerful AI safely even if it might be secretly working against its developers.
Worth knowing: Consults for governments and AI companies, including Google DeepMind and Anthropic.
Tool or datasetFor the curious
The Collective Intelligence Project · Weval · weval.org
Open platform where experts and communities write tests for AI models and publish the results, including checks on mental-health crisis responses and sycophancy.
Worth knowing: Scores are produced by AI 'judge' models, which can themselves make mistakes.