DMS
KV Cache Compression
Inference-Time Hyper-Scaling with KV Cache Compression
Inference-time scaling trades efficiency for increased reasoning accuracy by generating longer or more parallel sequences. However, in Transformer LLMs, generation cost is bottlenecked by the size...
https://arxiv.org/abs/2506.05345


Seonglae Cho