layer-by-layer inferenceairllmlyogavin • Updated 2026 Aug 23 10:40Qwen's new dense VL (Gated DeltaNet + Gated Attention, native vision) runs in 3.33GB of VRAM, measured end to end on one RTX 3090. Needs transformers 5.8+