AI Feature Dimensionality
Loading views...
Search

AI Feature Dimensionality

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2024 Apr 19 14:48
Editor
Editor
Seonglae ChoSeonglae Cho
Edited
Edited
2025 Feb 20 14:18
Refs
Refs
 
 
 
Toy Models of Superposition
It would be very convenient if the individual neurons of artificial neural networks corresponded to cleanly interpretable features of the input. For example, in an “ideal” ImageNet classifier, each neuron would fire only in the presence of a specific visual feature, such as the color red, a left-facing curve, or a dog snout. Empirically, in models we have studied, some of the neurons do cleanly map to features. But it isn't always the case that features correspond so cleanly to neurons, especially in large language models where it actually seems rare for neurons to correspond to clean features. This brings up many questions. Why is it that neurons sometimes align with features and sometimes don't? Why do some models and tasks have many of these clean neurons, while they're vanishingly rare in others?
Toy Models of Superposition
https://transformer-circuits.pub/2022/toy_model/index.html
 
 
 

Backlinks

AI ScientistFine Tuning DynamicsAI Memory CapacityMoE InterpretabilityReversing TransformerDropout

Recommendations

AI Feature Dimensionality
Copyright Seonglae Cho