AirLLM

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2024 Oct 21 20:44
Editor
Edited
Edited
2026 Aug 23 10:43
Refs
Refs

layer-by-layer inference

Qwen's new dense VL (Gated DeltaNet + Gated Attention, native vision) runs in 3.33GB of VRAM, measured end to end on one RTX 3090. Needs transformers 5.8+
 
 
 
 
 
 
 
 

Recommendations