memory bound 되서 좋다
Sliding-window beats linear attention
Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy. Every new token costs more than the previous one. For each additional token, the keys and...
https://arxiv.org/abs/2608.28444


Seonglae Cho