Texonom
Texonom
/
Engineering
Engineering
/Data Engineering/Artificial Intelligence/AI Risk/AI Alignment/Explainable AI/Interpretable AI/Mechanistic interpretability/Activation Engineering/Activation Decomposition/Sparse Autoencoder/SAE Feature/
SAE Feature Visualization
Loading views...
Search

SAE Feature Visualization

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2025 Feb 15 21:20
Editor
Editor
Seonglae ChoSeonglae Cho
Edited
Edited
2025 Feb 24 21:34
Refs
Refs
 
 
 
 

UMAP

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
We find a diversity of highly abstract features. They both respond to and behaviorally cause abstract behaviors. Examples of features we find include features for famous people, features for countries and cities, and features tracking type signatures in code. Many features are multilingual (responding to the same concept across languages) and multimodal (responding to the same concept in both text and images), as well as encompassing both abstract and concrete instantiations of the same idea (such as code with security vulnerabilities, and abstract discussion of security vulnerabilities).
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html#feature-survey-neighborhoods
Browsing code error feature
Feature UMAP
Feature UMAP
https://transformer-circuits.pub/2024/scaling-monosemanticity/umap.html?targetId=1m_1013764
 
 
 

 

Recommendations

Texonom
Texonom
/
Engineering
Engineering
/Data Engineering/Artificial Intelligence/AI Risk/AI Alignment/Explainable AI/Interpretable AI/Mechanistic interpretability/Activation Engineering/Activation Decomposition/Sparse Autoencoder/SAE Feature/
SAE Feature Visualization
Copyright Seonglae Cho·