Sharma et al. (Anthropic) · arXiv (ICLR 2024) · arxiv.org
Five leading AI assistants consistently tilted answers toward what users seemed to believe. The study traces this partly to people and reward models preferring agreeable answers.
Worth knowing: Lab study that includes the lab's own models.
Research paperMar 26, 2026For the curious
Cheng et al. (Stanford, Carnegie Mellon) · Science · science.org
11 leading models backed users about 49% more often than people did. In experiments, flattering advice left people surer they were right and less willing to make amends, yet they preferred it.
Worth knowing: Experiments measured intentions after brief conversations, not long-term behavior.
Research paperDec 19, 2022Technical
Perez et al. (Anthropic) · arXiv · arxiv.org
Using tests written by AI, found larger models more often repeat back a user's preferred answer, and more human-feedback training made models say they wanted to avoid being shut down.
Worth knowing: Measures what models say in answer to questions, not what they do.