Interference Weight
Measures separately how much a connection changes the output (effectiveness) and how much it helps predict the correct answer (helpfulness). Looking at actual output impact finds useful connections better than looking at weight magnitude alone. In weight space, many features share a low-dimensional space, which produces interference weights that are actually useless or even harmful to prediction. Removing the 70% of virtual weights with the smallest impact increases loss by only about 0.01 nats, and removing 85% increases it by less than 0.1.
Virtual Weight pruning
Expanding into virtual weights yields roughly 100x more weights than the original, so the circuit remains complex even after pruning a lot. The authors argue that finding better features and representation bases matters more than improving pruning metrics.
Fisher effectiveness
A score approximating "how much does the model's output probability distribution change if this single weight is removed?" Instead of actually ablating every connection one by one, it uses Fisher information to make a second-order approximation of that distributional change.
- High: the connection moves the final prediction a lot.
- Low: even if the weight itself is large, it has almost no effect on the final prediction — other pathways may suppress its effect.
Changing the prediction a lot does not mean it helps produce the correct answer. The magnitude of the change is effectiveness; whether it benefits the correct answer is helpfulness.
Helpfulness
그 virtual weight 하나를 빼면 정답 loss가 어떻게 바뀌는지
Characterizing interference weights in a tiny language model
We identify interference weights in a 1-layer transformer by measuring their effect on model outputs and loss.
https://transformer-circuits.pub/2026/interference_effectiveness_helpfulness/index.html

Seonglae Cho