SAE Feature Direction Loss
Tanh loss
Achieves Pareto-optimality by "minimizing feature activations while maintaining low output error"
Circuit Tracing: Revealing Computational Graphs in Language Models
We describe an approach to tracing the “step-by-step” computation involved when a model responds to a single prompt.
https://transformer-circuits.pub/2025/attribution-graphs/methods.html


Seonglae Cho