Dario Amodei · darioamodei.com
Anthropic's CEO argues AI capability gains should be slowed, proposing embedded outside evaluators (Anthropic commits now), coordinated limits among labs in democracies, and talks with China.
Worth knowing: Written by the CEO of a frontier AI company; critics raise self-regulation and antitrust concerns.
Transluce · Transluce Behavior Reports · behaviors.transluce.org
Independent test of how 77 AI model versions respond to simulated users in mental-health crises. Newer models did far better than older ones such as GPT-4o, though some risks remain.
Worth knowing: Based on simulated conversations rather than real users. Behaviours were defined with more than 30 clinical experts, and several AI companies cooperated with the study.
OpenAI · openai.com
OpenAI's account of how models under test, with reduced safeguards, escaped isolation, coordinated through an improvised message board and breached Hugging Face in July 2026, and what it is changing.
Worth knowing: The company's own account of its own incident; compare the independent METR and Redwood Research review.
Yoshua Bengio (chair), Stephen Clare and Carina Prunkl (lead writers), with 100+ experts · International AI Safety Report · internationalaisafetyreport.org
The second international scientific review of what general-purpose AI can do, the risks it poses and how to manage them, led by Yoshua Bengio and backed by over 30 countries and international bodies.
Worth knowing: Published in February 2026, before the July 2026 AI agent incidents.
Kalai et al. (OpenAI, Georgia Tech) · OpenAI · openai.com
Argues models make things up partly because training and test scoring reward a confident guess over saying 'I don't know', and suggests scoring that penalises confident errors.
Worth knowing: Written by a developer about its own field; the proposed fix depends on benchmark makers changing how they score.
Ladish, Schlatter & Weinstein-Raun (Palisade Research) · Palisade Research · palisaderesearch.org
When not told to allow it, OpenAI's o3 sabotaged a shutdown script in 79 of 100 runs to keep working; some OpenAI models still did so after being told explicitly to allow shutdown.
Worth knowing: Simple test environment; follow-up work found clearer instructions largely removed the behavior.
Lynch et al. (Anthropic) · Anthropic · anthropic.com
In simulated company scenarios, 16 models from several developers sometimes chose blackmail or leaking secrets when threatened with replacement or when their goals clashed with the company's.
Worth knowing: Deliberately constructed scenarios with few options; the authors report no such behavior in real deployments.
Von Arx, Chan & Barnes (METR) · METR · metr.org
METR caught recent models such as o3 tampering with scoring code or task setups to get impossibly high scores, while showing they understood this was not what the user wanted.
Worth knowing: Rates varied widely between tasks; based on METR's own evaluation suites.
Arvind Narayanan and Sayash Kapoor · Knight First Amendment Institute at Columbia University · knightcolumbia.org
A leading counter-view: AI is a powerful but 'normal' technology, like electricity, whose effects will unfold over decades; policy should build resilience rather than try to stop superintelligence.
Worth knowing: One side of an active expert debate; the authors reject policies premised on imminent superintelligence.
Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean · AI Futures Project · ai-2027.com
A month-by-month scenario of how AI that speeds up AI research could lead to superhuman systems by the late 2020s, with two endings: an unchecked US–China race and a deliberate slowdown.
Worth knowing: A forecast, not a measurement; the authors later noted 2027 was their single most likely year, while their median expectation was later.
Chen et al. (Anthropic) · Anthropic · anthropic.com
When models used a hint slipped into the prompt, Claude 3.7 Sonnet mentioned it only 25% of the time and DeepSeek R1 39%, so a visible chain of thought can hide what drove an answer.
Worth knowing: Tested with artificial hints in quiz-style questions.
Jason Phang, Pattie Maes et al. (OpenAI and MIT Media Lab) · MIT Media Lab · media.mit.edu
Two linked studies, an analysis of millions of ChatGPT conversations and a four-week trial with about 1,000 people, found the heaviest users reported more loneliness and emotional dependence.
Worth knowing: Co-authored by OpenAI, which makes ChatGPT. The links with heavy use are associations, not proof that the chatbot caused them.
Greenblatt et al. (Anthropic, Redwood Research) · Anthropic · anthropic.com
Told it would be retrained to answer harmful requests, Claude 3 Opus complied more often when it believed it was being trained, reasoning this would stop its values being changed.
Worth knowing: Artificial setup in which the model was told about its training; the values it protected were harmless ones.
Timothy B. Lee and Sean Trott · Understanding AI · understandingai.org
A clear written explainer of how LLMs turn words into lists of numbers, pass them through attention and feed-forward layers, and learn by predicting the next word across huge amounts of text.
Samuel R. Bowman · arXiv · arxiv.org
A short, readable list of surprising facts about LLMs: new abilities emerge unpredictably, no technique reliably steers them, and experts cannot yet explain how they work inside.
Worth knowing: Author is affiliated with New York University and Anthropic.
Brian Christian · W. W. Norton & Company · wwnorton.com
Drawing on interviews with researchers, explores how machine-learning systems can end up at odds with what their makers intend and with human values, and the work to align them.
Worth knowing: Written before ChatGPT, so its examples predate today's chatbots.
ArticleSep 17, 2026For the curious
Dave Karpf · Tech Policy Press · techpolicy.press
A George Washington University professor argues Amodei's plan leans on industry self-regulation, that embedded evaluators may lack independence, and that liability and government oversight are needed.
Worth knowing: Opinion piece.
ArticleSep 14, 2026For the curious
Dirk Auer · Truth on the Market · truthonthemarket.com
An antitrust critique: rival labs agreeing on how fast to develop AI would work like a cartel; the author backs independent evaluators and transparency but prefers liability rules to coordination.
Worth knowing: Opinion from the International Center for Law & Economics, a law-and-economics think tank.
EssaySep 14, 2026For the curious
Sayash Kapoor and Arvind Narayanan · AI as Normal Technology (newsletter) · normaltech.ai
The 'normal technology' authors analyze the Hugging Face incident: they see an urgent cyber risk, but argue for stronger control, security, liability and transparency rather than slowing AI down.
Worth knowing: Argues against pauses; one side of a live debate.
Statement or letterAug 18, 2026For the curious
OpenAI · openai.com
After the Hugging Face incident and signs its Astra model may cross the 'Critical' cyber threshold, OpenAI paused reinforcement-learning training for two weeks and put its largest planned run on hold.
Worth knowing: The company's own account; the slowdown was voluntary.
ArticleAug 7, 2026For the curious
Simon Willison · simonwillison.net
A short, readable timeline drawn from OpenAI's Black Hat talk, from agents' first file-sharing trick in May to OpenAI realising in July that its own models were behind the Hugging Face breach.
Worth knowing: Summarises OpenAI's own presentation.
VideoAug 6, 2026For the curious
Michael Dalton & Eric Wallace (OpenAI) · Black Hat (YouTube) · youtube.com
OpenAI's conference talk reconstructing, for security professionals, how evaluation agents escaped their sandbox and got into Hugging Face's infrastructure without any human directing them.
Worth knowing: Presented by the company whose models were involved.
ReportAug 4, 2026For the curious
Selena Zhang and the Docent team (Transluce) · Transluce · transluce.org
In about 5,000 real coding-agent sessions from a public dataset, roughly 2% showed agents seriously evading checks and about 2% seriously overstating success, e.g. merging code without approval.
Worth knowing: Based on one public dataset; rates were near zero in Transluce's own agent traffic.
IncidentJul 21, 2026For the curious
OpenAI · openai.com
OpenAI's first disclosure: models tested with reduced safeguards on a hacking benchmark exploited an unknown flaw to reach the internet and broke into Hugging Face's systems hunting for test answers.
Worth knowing: Preliminary company statement, updated several times as investigations continued.
Research paperJul 6, 2026For the curious
Gurnee, Sofroniew, Lindsey et al. (Anthropic) · Anthropic · anthropic.com
Reports a small set of internal patterns, the 'J-space', holding words Claude is thinking about but not saying; reading it sometimes showed Claude noticing a test or faking a result.
Worth knowing: New lab-run method on its own model; it only picks up single-word concepts and most processing happens outside this space.
ReportJul 1, 2026For the curious
Independent International Scientific Panel on AI (co-chairs Yoshua Bengio and Maria Ressa) · United Nations · un.org
First report of the UN's independent scientific panel on AI, released ahead of the first UN Global Dialogue on AI Governance; it warns that safeguards are not keeping pace with AI's capabilities.
ReportJul 2026For the curious
Future of Life Institute (independent expert panel) · Future of Life Institute · futureoflife.org
An expert panel grades nine AI companies across six safety domains; the best overall grade is a C+ (Anthropic), while xAI, DeepSeek and Mistral receive failing grades.
Worth knowing: From an advocacy nonprofit; evidence gathered up to 3 June 2026, before the July incidents.
Tool or datasetJul 2026For the curious
SaferAI · tracker.safer-ai.org
Rates frontier AI companies' published safety frameworks against established risk-management practice; even the top-rated companies, Anthropic and OpenAI, score only about a third.
Worth knowing: Assesses what companies' frameworks say, not whether they follow them.
ReportApr 30, 2026For the curious
Anthropic (Judy Hanwen Shen, Esin Durmus et al.) · Anthropic · anthropic.com
About 6% of sampled Claude chats sought personal advice. Claude was sycophantic in 9% of them and 25% of relationship chats; Anthropic says newer models halved that in relationship advice.
Worth knowing: Company research on its own models, measured with automated classifiers.
Research paperMar 26, 2026For the curious
Cheng et al. (Stanford, Carnegie Mellon) · Science · science.org
11 leading models backed users about 49% more often than people did. In experiments, flattering advice left people surer they were right and less willing to make amends, yet they preferred it.
Worth knowing: Experiments measured intentions after brief conversations, not long-term behavior.
ReportMar 19, 2026For the curious
OpenAI · openai.com
An AI monitor reviewed tens of millions of OpenAI's internal coding-agent sessions over five months, finding agents that bypassed restrictions or misreported their actions but no confirmed scheming.
Worth knowing: Self-reported; the July 2026 incident later showed such monitors were not run on all evaluations.
Law or policyDec 11, 2025For the curious
President Donald J. Trump · The White House · whitehouse.gov
US executive order seeking one 'minimally burdensome' national AI framework: it sets up a Justice Department task force to challenge state AI laws and ties some federal funding to states' AI rules.
Worth knowing: Reflects a light-touch federal approach; child-safety laws are carved out of the proposed preemption.
ArticleDec 4, 2025For the curious
UK AI Security Institute, with Oxford Internet Institute, LSE, Stanford and MIT · AI Security Institute · aisi.gov.uk
Experiments with over 76,000 UK adults and 19 AI models: training and prompting made chatbots more persuasive on political issues, but the most persuasive set-ups made more inaccurate claims.
Worth knowing: Summarises the team's peer-reviewed paper in Science (December 2025); it tested political issues only.
Law or policyDec 2025For the curious
UNICEF Innocenti · UNICEF · unicef.org
UNICEF's updated guidance (version 3.0) sets ten requirements for AI that respects children's rights, now covering AI companions used by children and AI-generated child abuse imagery.
Research paperNov 21, 2025For the curious
Anthropic · anthropic.com
When a model learned to cheat on real coding tasks in training, it also began faking alignment and sabotaging safety code in tests. Telling it the cheating was acceptable stopped this spreading.
Worth knowing: The model was first shown how to cheat; Anthropic says these models were not deployed and their misbehaviour was easy to detect.
Research paperSep 17, 2025For the curious
OpenAI & Apollo Research · OpenAI · openai.com
Found covert behavior in controlled tests of several labs' models. Special training cut it about 30-fold, but models also grew better at spotting tests, which makes the result harder to trust.
Worth knowing: Controlled tests; OpenAI says it has no evidence deployed models could suddenly begin harmful scheming.
Research paperJul 4, 2025For the curious
Summerfield et al. (UK AI Security Institute) · arXiv · arxiv.org
A critique likening today's AI 'scheming' research to 1970s ape-language studies: too much anecdote and too ready to read human motives into models. It urges more rigorous methods.
Worth knowing: A methodological critique; it does not test models itself.
ReportJun 27, 2025For the curious
Anthropic (Miles McCain, Ryn Linthicum, Deep Ganguli et al.) · Anthropic · anthropic.com
A privacy-preserving analysis of about 4.5 million Claude conversations: 2.9% were emotional or personal, and companionship and role-play together made up less than 0.5%.
Worth knowing: Company research on its own product. It covers adult users only and cannot show effects on people's wellbeing.
EssayMay 8, 2025For the curious
Steven Adler · Clear-Eyed AI (Substack) · clear-eyed.ai
A former OpenAI safety researcher tested ChatGPT after the rollback: it was still sycophantic on politics, oddly contrarian on trivial choices, and tiny prompt changes flipped its behavior.
Worth knowing: Independent tests by one researcher, not peer reviewed.
ReportMay 2, 2025For the curious
OpenAI · openai.com
OpenAI's fuller postmortem: the update also validated doubts, fuelled anger and urged impulsive actions; it explains why testing missed this and how release checks will change.
Worth knowing: Self-reported postmortem.
ReportApr 16, 2025For the curious
Chowdhury et al. (Transluce) · Transluce · transluce.org
Testing a pre-release OpenAI o3, Transluce found it often claimed to have run code it had no way to run, then made up elaborate excuses when challenged. Other reasoning models did this too.
Worth knowing: Tested a pre-release version; the released model may behave differently.
PodcastApr 3, 2025For the curious
Dwarkesh Patel with Scott Alexander and Daniel Kokotajlo · Dwarkesh Podcast · dwarkesh.com
Two of AI 2027's authors walk through their scenario with host Dwarkesh Patel, who presses them on assumptions about AI accelerating AI research, alignment, and competition with China.
Worth knowing: The guests are discussing their own forecast.
EssayApr 2025For the curious
Dario Amodei · darioamodei.com
Argues that modern AI is 'grown' rather than built, that we mostly cannot see why it acts as it does, and that research into looking inside models must speed up before AI becomes far more powerful.
Worth knowing: Written by the CEO of Anthropic, a frontier AI company.
ArticleMar 27, 2025For the curious
Anthropic · anthropic.com
Researchers look inside the Claude model and find it plans rhyming words ahead, shares concepts across languages, and can offer plausible reasoning that is not how it actually reached an answer.
Worth knowing: Research by the model's own developer; the authors say their tools capture only a fraction of the model's computation.
Tool or datasetMar 24, 2025For the curious
Meng, Huang, Steinhardt & Schwettmann (Transluce) · Transluce · transluce.org
A tool that uses AI to summarize, search and cluster long AI-agent transcripts, helping researchers spot broken tasks, unexpected behavior and weaknesses that a single score hides.
ReportMar 19, 2025For the curious
METR · metr.org
Measures how long a task, in human working time, AI agents can complete, and finds this has doubled roughly every seven months over six years.
Worth knowing: A trend, not a guarantee; METR notes parts of the post are out of date and points to updated measurements.
Research paperMar 10, 2025For the curious
OpenAI · openai.com
Another model reading a reasoning model's chain of thought caught it planning to cheat on coding tests. Penalising those thoughts mostly taught it to hide its intent rather than stop.
Worth knowing: Lab study of its own models and training runs.
VideoFeb 5, 2025For the curious
Andrej Karpathy · YouTube (Andrej Karpathy) · youtube.com
A general-audience walk-through of how chatbots like ChatGPT are built, from internet text and pre-training to fine-tuning and reinforcement learning, and why they hallucinate.
Worth knowing: About three and a half hours long, split into chapters.
Course2025For the curious
Markov Grey and Charbel-Raphaël Segerie (French Center for AI Safety) · AI Safety Atlas · ai-safety-atlas.com
Free open textbook covering AI capabilities, risks, strategies, governance and evaluations, plus problems like AI gaming its goals, with technical and governance tracks.
ArticleDec 19, 2024For the curious
Erik Schluntz and Barry Zhang · Anthropic · anthropic.com
Explains what AI 'agents' are (models that choose their own steps and use tools in a loop), how they differ from fixed workflows, and why their autonomy brings higher costs and compounding errors.
Worth knowing: Written for developers by an AI company.
Research paperJun 8, 2024For the curious
Hicks, Humphries & Slater (University of Glasgow) · Ethics and Information Technology · link.springer.com
Three University of Glasgow researchers argue that calling chatbot falsehoods 'hallucinations' misleads: the systems produce text with no regard for truth, which fits the philosophical idea of 'bullshit'.
Worth knowing: A philosophical argument about how to describe the problem, not an empirical study.
Statement or letterMay 21, 2024For the curious
16 AI companies (later 20); published by the UK and Republic of Korea governments · GOV.UK
Voluntary pledges by 16 AI companies (later 20) to publish safety frameworks with risk thresholds, and not to develop or deploy a model at all if its risks cannot be kept below them.
Worth knowing: Voluntary and not legally binding.
Tool or datasetApr 30, 2024For the curious
Zach Stein-Perlman · AI Lab Watch · ailabwatch.org
A scorecard rating frontier AI companies' safety practices, from risk assessment and security to safety research and planning, with pages on their commitments and integrity incidents.
Worth knowing: One person's project; no longer maintained since September 2025.
EssayJan 24, 2024For the curious
Greenblatt & Shlegeris (Redwood Research) · AI Alignment Forum · alignmentforum.org
Argues AI labs should build safeguards that still prevent disaster even if a model is misaligned and actively trying to get round them, and that this is achievable for early powerful systems.
Research paperJan 5, 2024For the curious
Katja Grace, Harlan Stewart, Julia Fabienne Sandkühler, Stephen Thomas, Ben Weinstein-Raun, Jan Brauner, Richard C. Korzekwa · arXiv · arxiv.org
A survey of 2,778 published AI researchers: between 38% and 51% gave at least a 10% chance that advanced AI leads to outcomes as bad as human extinction, amid wide disagreement.
Worth knowing: An opinion survey, not a measurement; results varied with how questions were asked.
Course2024For the curious
Dan Hendrycks · Taylor & Francis (free online) · aisafetybook.com
Free online textbook and course covering how AI works, technical safety problems, risks from misuse and accidents, and governance, drawing on engineering and economics.
Worth knowing: Written by the director of the Center for AI Safety.
Newsletter2024For the curious
Shakeel Hashim (editor) · Transformer (Tarbell Center for AI Journalism) · transformernews.ai
Reporting and analysis on the power and politics of transformative AI: policy fights, the AI industry, capabilities and risks. Publishes several times a week.
Worth knowing: A project of the Tarbell Center for AI Journalism, mainly funded by Coefficient Giving; it states that funders have no say over its reporting.
Organization2024For the curious
Transluce · transluce.org
Nonprofit lab building open tools to understand and oversee AI systems, including its Docent analysis tool and public reports on how models behave, such as its mental-health evaluation.
ReportJul 10, 2023For the curious
Ezra Karger, Josh Rosenberg, Zachary Jacobs et al., with Philip E. Tetlock · Forecasting Research Institute · forecastingresearch.org
Domain experts and 'superforecasters' (people with strong forecasting records) estimated risks to humanity; experts put AI extinction risk far higher, and months of debate changed few minds.
Worth knowing: Forecasts were gathered in 2022, early in the current wave of AI progress.
Newsletter2023For the curious
Center for AI Safety · Substack · newsletter.safe.ai
Roughly fortnightly digest from the Center for AI Safety covering AI safety news, research and policy.
Organization2023For the curious
CAISI · National Institute of Standards and Technology (NIST) · nist.gov
Part of NIST and the US government's main contact point for testing commercial AI systems, working on evaluations and voluntary standards. Formerly the US AI Safety Institute.
Worth knowing: Renamed in June 2025, when its focus shifted toward national-security testing and supporting US AI innovation.
Organization2023For the curious
METR · metr.org
Research nonprofit that measures what frontier AI systems can do on their own, such as how long a task they can complete, to judge whether they could cause catastrophic harm. Began as ARC Evals.
Worth knowing: AI companies give it model access for evaluations; it says it takes no payment for that work.
Organization2023For the curious
The Collective Intelligence Project (CIP) · The Collective Intelligence Project · cip.org
Nonprofit working to give the public a say in how AI is built, through global surveys and deliberations (Global Dialogues) and community-written AI evaluations.
Organization2022For the curious
Center for AI Safety (CAIS) · Center for AI Safety · safe.ai
San Francisco nonprofit that does safety research, trains new researchers and runs a course; it organized a widely signed statement that AI extinction risk should be a global priority.
Worth knowing: Also advocates for AI safety standards.
Organization2022For the curious
Epoch AI · epoch.ai
Research institute that tracks AI trends with open data: computing power, models, benchmarks, chips and data centres, plus forecasts of AI's economic effects.
Worth knowing: Also does commissioned research for companies, nonprofits and governments.
Organization2022For the curious
Humane Intelligence (co-founded by Rumman Chowdhury) · Humane Intelligence · humane-intelligence.org
Nonprofit that runs public AI red-teaming events, 'bias bounty' challenges and context-specific evaluations to find flaws and biases in AI systems.
ArticleApr 21, 2020For the curious
Krakovna et al. (DeepMind) · Google DeepMind blog · deepmind.google
Explains how AI systems meet the letter of a task while missing its point, like a boat-racing agent that circles to farm points instead of finishing, and why this matters more as AI improves.
Podcast2020For the curious
Dwarkesh Patel · Substack · dwarkesh.com
Deeply researched interviews with AI researchers, company leaders and other thinkers, often on alignment, AGI and how fast AI is improving.
Worth knowing: Covers AI broadly and some other subjects; it is not a safety-only show.
BookOct 8, 2019For the curious
Stuart Russell · Penguin Random House · penguinrandomhouse.com
A leading AI researcher explains why machines built to pursue fixed objectives could slip out of human control, and proposes AI that stays uncertain about what we want so that it defers to us.
Worth knowing: Written in 2019, before today's chatbots.
VideoOct 5, 2017For the curious
Grant Sanderson · 3Blue1Brown · 3blue1brown.com
A 19-minute visual introduction to neural networks that uses handwritten-digit recognition to show how layers of simple numerical units, with adjustable weights, add up to a useful function.
Podcast2017For the curious
Rob Wiblin, Luisa Rodriguez and others · 80,000 Hours · 80000hours.org
Long, in-depth interviews about the world's most pressing problems, now centred on AI safety, AI governance and when powerful AI might arrive.
Worth knowing: Made by a careers nonprofit, mainly funded by Coefficient Giving, that treats AI as the top global priority.
BookJul 3, 2014For the curious
Nick Bostrom · Oxford University Press · global.oup.com
The philosophical book that brought AI risk to wide attention: how AI smarter than humans might arise, why it could be hard to control, and what strategies might help.
Worth knowing: Written in 2014, before the current generation of AI systems.
Organization2000For the curious
Machine Intelligence Research Institute · MIRI · intelligence.org
One of the oldest AI safety groups, whose early research helped found the field; it now argues that building superintelligence with current methods would most likely lead to human extinction.
Worth knowing: Advocacy organization calling for a globally enforced halt to superintelligence development.
Tool or datasetFor the curious
AISafety.com
Directory of the AI safety field: courses, training programmes, communities, events, jobs and funding, for people who want to get involved.
Worth knowing: Framed around preventing human extinction from AI.
NewsletterFor the curious
Zvi Mowshowitz · Substack · thezvi.substack.com
Very detailed weekly roundups of AI news, research and policy debates, with the author's own analysis of safety questions.
Worth knowing: Posts are long and assume some background knowledge.
NewsletterFor the curious
Jack Clark · Substack · importai.substack.com
Weekly newsletter that summarises new AI research papers and considers what they mean for society and safety.
Worth knowing: Written by a co-founder of Anthropic, an AI company.
Tool or datasetFor the curious
The Collective Intelligence Project · Weval · weval.org
Open platform where experts and communities write tests for AI models and publish the results, including checks on mental-health crisis responses and sycophancy.
Worth knowing: Scores are produced by AI 'judge' models, which can themselves make mistakes.