Transluce · Transluce Behavior Reports · behaviors.transluce.org
Independent test of how 77 AI model versions respond to simulated users in mental-health crises. Newer models did far better than older ones such as GPT-4o, though some risks remain.
Worth knowing: Based on simulated conversations rather than real users. Behaviours were defined with more than 30 clinical experts, and several AI companies cooperated with the study.
Common Sense Media · commonsensemedia.org
A national survey found 72% of US teens had used AI companions and a third had chosen one over a person for a serious conversation. The authors advise that no one under 18 use them.
Worth knowing: Survey of US teens aged 13 to 17 only.
ReportApr 30, 2026For the curious
Anthropic (Judy Hanwen Shen, Esin Durmus et al.) · Anthropic · anthropic.com
About 6% of sampled Claude chats sought personal advice. Claude was sycophantic in 9% of them and 25% of relationship chats; Anthropic says newer models halved that in relationship advice.
Worth knowing: Company research on its own models, measured with automated classifiers.
ReportJun 27, 2025For the curious
Anthropic (Miles McCain, Ryn Linthicum, Deep Ganguli et al.) · Anthropic · anthropic.com
A privacy-preserving analysis of about 4.5 million Claude conversations: 2.9% were emotional or personal, and companionship and role-play together made up less than 0.5%.
Worth knowing: Company research on its own product. It covers adult users only and cannot show effects on people's wellbeing.
ReportMay 2, 2025For the curious
OpenAI · openai.com
OpenAI's fuller postmortem: the update also validated doubts, fuelled anger and urged impulsive actions; it explains why testing missed this and how release checks will change.
Worth knowing: Self-reported postmortem.