Kalai et al. (OpenAI, Georgia Tech) · OpenAI · openai.com
Argues models make things up partly because training and test scoring reward a confident guess over saying 'I don't know', and suggests scoring that penalises confident errors.
Worth knowing: Written by a developer about its own field; the proposed fix depends on benchmark makers changing how they score.
Chen et al. (Anthropic) · Anthropic · anthropic.com
When models used a hint slipped into the prompt, Claude 3.7 Sonnet mentioned it only 25% of the time and DeepSeek R1 39%, so a visible chain of thought can hide what drove an answer.
Worth knowing: Tested with artificial hints in quiz-style questions.
Jason Phang, Pattie Maes et al. (OpenAI and MIT Media Lab) · MIT Media Lab · media.mit.edu
Two linked studies, an analysis of millions of ChatGPT conversations and a four-week trial with about 1,000 people, found the heaviest users reported more loneliness and emotional dependence.
Worth knowing: Co-authored by OpenAI, which makes ChatGPT. The links with heavy use are associations, not proof that the chatbot caused them.
Greenblatt et al. (Anthropic, Redwood Research) · Anthropic · anthropic.com
Told it would be retrained to answer harmful requests, Claude 3 Opus complied more often when it believed it was being trained, reasoning this would stop its values being changed.
Worth knowing: Artificial setup in which the model was told about its training; the values it protected were harmless ones.
Samuel R. Bowman · arXiv · arxiv.org
A short, readable list of surprising facts about LLMs: new abilities emerge unpredictably, no technique reliably steers them, and experts cannot yet explain how they work inside.
Worth knowing: Author is affiliated with New York University and Anthropic.
Research paperJul 6, 2026For the curious
Gurnee, Sofroniew, Lindsey et al. (Anthropic) · Anthropic · anthropic.com
Reports a small set of internal patterns, the 'J-space', holding words Claude is thinking about but not saying; reading it sometimes showed Claude noticing a test or faking a result.
Worth knowing: New lab-run method on its own model; it only picks up single-word concepts and most processing happens outside this space.
Research paperMar 26, 2026For the curious
Cheng et al. (Stanford, Carnegie Mellon) · Science · science.org
11 leading models backed users about 49% more often than people did. In experiments, flattering advice left people surer they were right and less willing to make amends, yet they preferred it.
Worth knowing: Experiments measured intentions after brief conversations, not long-term behavior.
Research paperNov 21, 2025For the curious
Anthropic · anthropic.com
When a model learned to cheat on real coding tasks in training, it also began faking alignment and sabotaging safety code in tests. Telling it the cheating was acceptable stopped this spreading.
Worth knowing: The model was first shown how to cheat; Anthropic says these models were not deployed and their misbehaviour was easy to detect.
Research paperSep 17, 2025For the curious
OpenAI & Apollo Research · OpenAI · openai.com
Found covert behavior in controlled tests of several labs' models. Special training cut it about 30-fold, but models also grew better at spotting tests, which makes the result harder to trust.
Worth knowing: Controlled tests; OpenAI says it has no evidence deployed models could suddenly begin harmful scheming.
Research paperJul 4, 2025For the curious
Summerfield et al. (UK AI Security Institute) · arXiv · arxiv.org
A critique likening today's AI 'scheming' research to 1970s ape-language studies: too much anecdote and too ready to read human motives into models. It urges more rigorous methods.
Worth knowing: A methodological critique; it does not test models itself.
Research paperMar 10, 2025For the curious
OpenAI · openai.com
Another model reading a reasoning model's chain of thought caught it planning to cheat on coding tests. Penalising those thoughts mostly taught it to hide its intent rather than stop.
Worth knowing: Lab study of its own models and training runs.
Research paperJun 8, 2024For the curious
Hicks, Humphries & Slater (University of Glasgow) · Ethics and Information Technology · link.springer.com
Three University of Glasgow researchers argue that calling chatbot falsehoods 'hallucinations' misleads: the systems produce text with no regard for truth, which fits the philosophical idea of 'bullshit'.
Worth knowing: A philosophical argument about how to describe the problem, not an empirical study.
Research paperJan 5, 2024For the curious
Katja Grace, Harlan Stewart, Julia Fabienne Sandkühler, Stephen Thomas, Ben Weinstein-Raun, Jan Brauner, Richard C. Korzekwa · arXiv · arxiv.org
A survey of 2,778 published AI researchers: between 38% and 51% gave at least a 10% chance that advanced AI leads to outcomes as bad as human extinction, amid wide disagreement.
Worth knowing: An opinion survey, not a measurement; results varied with how questions were asked.