Learning Latent Variable of sequential data by predicting a future data with Contrastive Learning.
Relative patch prediction
for easier decomposition?
Representational Simplicity and Circuit Size Dissociate in a...
Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model's computation is easier to reverse-engineer. Whether this representational...
https://arxiv.org/abs/2609.35890


Seonglae Cho