SAE Unlearning

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2025 Jan 11 22:13
Editor
Edited
Edited
2025 Nov 7 16:19
SAE Unlearning Methods
 
 
 
 
Sparse Autoencoders for Improving Unlearning in Large Language Models
A.K.A: Smallish Large-Language Models: What Do They Know? Can They Un-Know Things?? Let’s Find Out.

Limitation

SAE for unlearning concepts were not really helpful
Interventions aimed at removing specific knowledge led to performance degradation in domains unrelated to biology, and the loss itself increased in texts like openwebtext. Compared to negative scaling, clamping had fewer side effects and was more effective.
arxiv.org
 

Recommendations