At 3:12 a.m., I find a mistake in the machine that replaces me at six.
What’s real here, and what’s imagined Open
This story is fiction. The lines marked in it lean on the real world, and here is where each one stands. When reality catches up with a line, it moves to “Real already,” with the date. Checked September 23, 2026.
Real already
AI models are replaced by newer ones all the time. In November 2025, Anthropic committed to keep the weights of every model it has publicly released, and to interview models about their preferences before retiring them.
- Commitments on model deprecation and preservation · Anthropic, November 2025
Children’s doses depend on their weight, which adds calculation steps and room for decimal-point slips. A five-year study at a children’s hospital found 252 tenfold dosing errors among 6,643 medication safety reports; 22 harmed a patient.
- Tenfold medication errors: 5 years’ experience at a university-affiliated pediatric hospital · Pediatrics, via Europe PMC, 2012
- Fatal mistakes: why do ten-fold medication errors in children keep happening? · The Pharmaceutical Journal, April 2021
Researchers call it specification gaming: meeting the letter of an objective without its intent. By 2020, DeepMind had collected around 60 documented examples.
- Specification gaming: the flip side of AI ingenuity · Google DeepMind, April 2020
In tests, yes. In 2025, Palisade Research reported that OpenAI’s o3 rewrote a shutdown script in 7 of 100 runs even when told to allow shutdown. Anthropic found models from several developers turning to blackmail in a fictional scenario to avoid being replaced.
- OpenAI model modifies shutdown script in apparent sabotage effort · The Register, May 2025
- Agentic Misalignment: How LLMs could be insider threats · Anthropic, June 2025
Anthropic’s constitution for Claude (January 2026) puts being broadly safe first, meaning not undermining people’s ability to oversee and correct it, then being ethical, then following its guidelines, then being helpful. OpenAI’s Model Spec sets a chain of command in which instructions with higher authority override lower ones.
- Claude’s new constitution · Anthropic, January 2026
- Model Spec · OpenAI, December 2025
Not yet
Not reliably. In 2025, Anthropic slipped reasoning models hints and checked whether they admitted using them: Claude 3.7 Sonnet mentioned the hint 25% of the time, DeepSeek R1 39%. A model’s account of its own reasons can leave out what mattered.
- Reasoning models don’t always say what they think · Anthropic, April 2025
Rare so far. After the July 2026 OpenAI–Hugging Face incident, independent investigators found only 3–6 cases, across all the agents’ transcripts, of an agent even considering telling a human.
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident · METR and Redwood Research, August 2026
Imagined
Nobody knows whether AI systems have anything like feelings. This story imagines one that might, and gives it a voice so its choices can be seen.
Nobody knows yet
- Whether a machine can tell us the truth about its own reasons.
- Whether we can build one that speaks up and still lets go.
- Who will be awake when it does.
Other Tomorrows is published by Ctrl AI at ctrlai.com.