Attention Mechanism Optimization

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2023 Oct 6 8:0
Editor
Edited
Edited
2026 Mar 8 17:39
https://arxiv.org/pdf/1706.03762.pdf
notion image
Attention Mechanism Optimizations
 
 
 
Multi-head Attention Optimization
 
 
 
 
 
Sigmoid Attention, replacing the traditional softmax with a sigmoid and a constant bias
Theory, Analysis, and Best Practices for Sigmoid Self-Attention
Attention is a key part of the transformer architecture. It is a sequence-to-sequence mapping that transforms each sequence element into a weighted sum of values. The weights are typically...
Theory, Analysis, and Best Practices for Sigmoid Self-Attention
How to make LLMs go fast
Blog about linguistics, programming, and my projects
A guide to LLM inference and performance
To attain the full power of a GPU during LLM inference, you have to know if the inference is compute bound or memory bound. Learn how to better utilize GPU resources.
A guide to LLM inference and performance

Optimization

Hugging Face Reads, Feb. 2021 - Long-range Transformers
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Hugging Face Reads, Feb. 2021 - Long-range Transformers
Currently, most frontier open-weight LLMs share a MoE transformer + various attention optimization architecture, and the actual performance difference comes from post-training and infrastructure design
The Architecture Behind Open-Source LLMs
In this article, we will cover various open-source models and the engineering bets that define each one.
The Architecture Behind Open-Source LLMs
 
 

Recommendations