TFA

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2026 Jan 3 22:33
Editor
Edited
Edited
2026 Apr 23 19:16
Refs

Temporal Feature Analysis

  • (predictable / slow-moving / context) the predictable component from past context
  • (novel / fast-moving / residual) new information (residual) not explained by context
TFA creates a direction that explains the current using past activations , implementing this in an attention form such as
NEPA
. Novel component: "apply SAE to the residual"
 
 
 

Token sequence

SAE Training Dataset Influence in Feature Matching and a Hypothesis on Position Features — LessWrong
Abstract Sparse Autoencoders (SAEs) linearly extract interpretable features from a large language model's intermediate representations. However, the…
SAE Training Dataset Influence in Feature Matching and a Hypothesis on Position Features — LessWrong
 
 
 

Recommendations