Texonom
Texonom
/
Engineering
Engineering
/Data Engineering/Artificial Intelligence/Machine Learning/Neural Network/Neural Network Structure/Seq2Seq/Attention Mechanism/Reversing Transformer/
Transformer Layer Depth
Loading views...
Search

Transformer Layer Depth

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2025 Dec 22 18:25
Editor
Editor
Seonglae ChoSeonglae Cho
Edited
Edited
2025 Dec 22 23:31
Refs
Refs
AI Feature Composition
Layer-wise changes: early=context structure, middle=semantics, late=output control
 
 
 
Distributed Representations: Composition & Superposition
Distributed representations are a classic idea in both neuroscience and connectionist approaches to AI. We're often asked how our work on superposition relates to it. Since publishing our original paper on superposition, we've had more time to reflect on the relationship between the topics and discuss it with people, and wanted to expand on our earlier discussion in the related work section and share a few thoughts. (We care a lot about superposition and the structure of distributed representations because decomposing representations into independent components is necessary to escape the curse of dimensionality and understand neural networks.)
Distributed Representations: Composition & Superposition
https://transformer-circuits.pub/2023/superposition-composition/index.html
 
 
 

Backlinks

AI Feature

Recommendations

Texonom
Texonom
/
Engineering
Engineering
/Data Engineering/Artificial Intelligence/Machine Learning/Neural Network/Neural Network Structure/Seq2Seq/Attention Mechanism/Reversing Transformer/
Transformer Layer Depth
Copyright Seonglae Cho·Sign in