Texonom
Texonom
/
Engineering
Engineering
/Data Engineering/Artificial Intelligence/AI Risk/AI Alignment/Explainable AI/Interpretable AI/Mechanistic interpretability/Activation Engineering/Activation Decomposition/Sparse Autoencoder/
Tied SAE
Loading views...
Search

Tied SAE

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2025 Feb 25 12:14
Editor
Editor
Seonglae ChoSeonglae Cho
Edited
Edited
2025 Mar 13 20:4
Refs
Refs
cosine similarity loss between feature direction of encoder and decoder matrix
 
 
 
 
 

Untied SAE

Toy Models of Feature Absorption in SAEs — LessWrong
TLDR; In previous work, we found a problematic form of feature splitting called "feature absorption" when analyzing Gemma Scope SAEs. We hypothesized…
Toy Models of Feature Absorption in SAEs — LessWrong
https://www.lesswrong.com/posts/kcg58WhRxFA9hv9vN/toy-models-of-feature-absorption-in-saes
Toy Models of Feature Absorption in SAEs — LessWrong
 
 

Recommendations

Texonom
Texonom
/
Engineering
Engineering
/Data Engineering/Artificial Intelligence/AI Risk/AI Alignment/Explainable AI/Interpretable AI/Mechanistic interpretability/Activation Engineering/Activation Decomposition/Sparse Autoencoder/
Tied SAE
Copyright Seonglae Cho·