Transluce · Transluce Behavior Reports · behaviors.transluce.org
Independent test of how 77 AI model versions respond to simulated users in mental-health crises. Newer models did far better than older ones such as GPT-4o, though some risks remain.
Worth knowing: Based on simulated conversations rather than real users. Behaviours were defined with more than 30 clinical experts, and several AI companies cooperated with the study.
Jason Phang, Pattie Maes et al. (OpenAI and MIT Media Lab) · MIT Media Lab · media.mit.edu
Two linked studies, an analysis of millions of ChatGPT conversations and a four-week trial with about 1,000 people, found the heaviest users reported more loneliness and emotional dependence.
Worth knowing: Co-authored by OpenAI, which makes ChatGPT. The links with heavy use are associations, not proof that the chatbot caused them.
ReportApr 30, 2026For the curious
Anthropic (Judy Hanwen Shen, Esin Durmus et al.) · Anthropic · anthropic.com
About 6% of sampled Claude chats sought personal advice. Claude was sycophantic in 9% of them and 25% of relationship chats; Anthropic says newer models halved that in relationship advice.
Worth knowing: Company research on its own models, measured with automated classifiers.
Research paperMar 26, 2026For the curious
Cheng et al. (Stanford, Carnegie Mellon) · Science · science.org
11 leading models backed users about 49% more often than people did. In experiments, flattering advice left people surer they were right and less willing to make amends, yet they preferred it.
Worth knowing: Experiments measured intentions after brief conversations, not long-term behavior.
ArticleDec 4, 2025For the curious
UK AI Security Institute, with Oxford Internet Institute, LSE, Stanford and MIT · AI Security Institute · aisi.gov.uk
Experiments with over 76,000 UK adults and 19 AI models: training and prompting made chatbots more persuasive on political issues, but the most persuasive set-ups made more inaccurate claims.
Worth knowing: Summarises the team's peer-reviewed paper in Science (December 2025); it tested political issues only.
Law or policyDec 2025For the curious
UNICEF Innocenti · UNICEF · unicef.org
UNICEF's updated guidance (version 3.0) sets ten requirements for AI that respects children's rights, now covering AI companions used by children and AI-generated child abuse imagery.
ReportJun 27, 2025For the curious
Anthropic (Miles McCain, Ryn Linthicum, Deep Ganguli et al.) · Anthropic · anthropic.com
A privacy-preserving analysis of about 4.5 million Claude conversations: 2.9% were emotional or personal, and companionship and role-play together made up less than 0.5%.
Worth knowing: Company research on its own product. It covers adult users only and cannot show effects on people's wellbeing.
ReportMay 2, 2025For the curious
OpenAI · openai.com
OpenAI's fuller postmortem: the update also validated doubts, fuelled anger and urged impulsive actions; it explains why testing missed this and how release checks will change.
Worth knowing: Self-reported postmortem.
Organization2024For the curious
Transluce · transluce.org
Nonprofit lab building open tools to understand and oversee AI systems, including its Docent analysis tool and public reports on how models behave, such as its mental-health evaluation.
Organization2023For the curious
The Collective Intelligence Project (CIP) · The Collective Intelligence Project · cip.org
Nonprofit working to give the public a say in how AI is built, through global surveys and deliberations (Global Dialogues) and community-written AI evaluations.
Organization2022For the curious
Humane Intelligence (co-founded by Rumman Chowdhury) · Humane Intelligence · humane-intelligence.org
Nonprofit that runs public AI red-teaming events, 'bias bounty' challenges and context-specific evaluations to find flaws and biases in AI systems.
Tool or datasetFor the curious
The Collective Intelligence Project · Weval · weval.org
Open platform where experts and communities write tests for AI models and publish the results, including checks on mental-health crisis responses and sycophancy.
Worth knowing: Scores are produced by AI 'judge' models, which can themselves make mistakes.