Influence Function

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2025 Jun 17 9:44
Editor
Edited
Edited
2025 Jun 17 10:52
Refs

Calculate the changes in logits caused by data points

influence functions can identify problematic documents within the continued pretraining corpus, enabling more targeted curation of safer language-specific data. There is currently very limited work on analyzing training-example-to-output relationships for multilingual safety-relevant behaviors.

EK-FAC

Since the model parameters are in the billions, it is physically impossible to directly calculate, store, and invert the entire Hessian matrix. Therefore, we approximate it using Kronecker factorization to represent the full layer Hessian as the Kronecker product of these two matrices. Additionally, instead of using simple gradient information, we use second-order information from the
Fisher Information Matrix
for more accurate estimation.
I have a question about the naming convention. Currently we are using names like Agent_1 and Agent_0, but Raj mentioned that clients might prefer more readable names such as Search Agent or SNS Writer Agent.

Results

Rare tokens tend to have high influence from a small number of training examples, while frequent tokens show more distributed influence. As model size increases, even with less token overlap, semantically similar sequences show high influence, demonstrating higher levels of generalization and abstraction. While strong generalization was observed in high-resource languages, low-resource languages showed limited performance. Additionally, exact matching datasets were important in mathematical programming. When word order in sentences is reversed, the influence almost disappears, confirming that LLMs heavily utilize sequential information.

Limitation

In non-convex optimization regions, Hessian approximation accuracy may decrease, and certain phenomena like order sensitivity cannot be fully explained. Future work requires more sophisticated curvature approximation methods and expansion of candidate filtering techniques.
 
 

EK-FAC & Eigenvalue Correction with TF-IDF filtering
Nelson Elhage

arxiv.org
 
 
 

Backlinks

Recommendations