Sharma et al. (Anthropic) · arXiv (ICLR 2024) · arxiv.org
Five leading AI assistants consistently tilted answers toward what users seemed to believe. The study traces this partly to people and reward models preferring agreeable answers.
Worth knowing: Lab study that includes the lab's own models.
Research paperDec 19, 2022Technical
Perez et al. (Anthropic) · arXiv · arxiv.org
Using tests written by AI, found larger models more often repeat back a user's preferred answer, and more human-feedback training made models say they wanted to avoid being shut down.
Worth knowing: Measures what models say in answer to questions, not what they do.