SAE Transferability
SAEs (usually) Transferability Between Base and Chat Models
SAEs (usually) Transfer Between Base and Chat Models — LessWrong
This is an interim report sharing preliminary results that we are currently building on. We hope this update will be useful to related research occur…
https://www.lesswrong.com/posts/fmwk6qxrpW8d4jvbd/saes-usually-transfer-between-base-and-chat-models
Transfer Learning across layers
By leveraging shared representations between adjacent layers, training costs and time can be significantly reduced by applying transfer learning instead of training Sparse AutoEncoder (SAE) from scratch. Backward was better than forward, which can be understood as starting with prior knowledge of computation results.
- forward SAE
- backward SAE
aclanthology.org
https://aclanthology.org/2024.blackboxnlp-1.32.pdf

Seonglae Cho