Negative Results for Sparse Autoencoders On Downstream Tasks and Deprioritising SAE Research (Mechanistic Interpretability Team Progress Update)
Smith, Rajamanoharan et al. (Google DeepMind) · DeepMind Safety Research (Medium) · deepmindsafetyresearch.medium.com
Google DeepMind found sparse autoencoders, a popular tool for finding concepts inside models, did worse than simple probes at detecting harmful intent, and scaled back its work on them.
Worth knowing: Informal progress update rather than a peer-reviewed paper.