Phase Transition
What derives Phase change
First, the window where the phase change happens doesn’t appear to correspond to a scheduled change in learning rate, warmup, or weight decay; there is not some known exogenous factor precipitating everything. Second, we tried out training some of the small models on a different dataset, and we observed the phase change develop in the same way.
Is there a way we could understand what "fraction of a dimension" a specific feature gets?
- 40ff6a0 - 1-hop rebac-based authz on asset
- 38e2751 - 1-hop rebac-based authz on artifact
- 38d83ae - editable artifact/asset authz
- 8dfaf7f - task asset/artifact authority passing
Reverse engineering with Induction head in Multi-head Attention
In-context Learning and Induction Heads
As Transformer generative models continue to scale and gain increasing real world use , addressing their associated safety problems becomes increasingly important. Mechanistic interpretability – attempting to reverse engineer the detailed computations performed by the model – offers one possible avenue for addressing these safety issues. If we can understand the internal structures that cause Transformer models to produce the outputs they do, then we may be able to address current safety problems more systematically, as well as anticipating safety problems in future more powerful models. Note that mechanistic interpretability is a subset of the broader field of interpretability, which encompasses many different methods for explaining the outputs of a neural network. Mechanistic interpretability is distinguished by a specific focus on trying to systematically characterize the internal circuitry of a neural net.
https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html
A Mechanistic Interpretability Analysis of Grokking — LessWrong
A significantly updated version of this work is now on Arxiv and was published as a spotlight paper at ICLR 2023 …
https://www.lesswrong.com/posts/N6WM6hs7RQMKDhYjB/a-mechanistic-interpretability-analysis-of-grokking
phase transition
Towards Developmental Interpretability — LessWrong
Developmental interpretability is a research agenda that has grown out of a meeting of the Singular Learning Theory (SLT) and AI alignment communitie…
https://www.lesswrong.com/posts/TjaeCWvLZtEDAS5Ex/towards-developmental-interpretability

Seonglae Cho