In terms of information entropy, tokens with lower perplexity (PPL) contribute less to the overall entropy gains of the language model. In other words, removing tokens with lower perplexity has a relatively minor impact on the LLM’s comprehension of the context.
2 LLMLinguamicrosoft • Updated 2026 Sep 2 14:3
LLMLingua
microsoft • Updated 2026 Sep 2 14:3
LLMLingua-2: Data Distillation for Efficient and Faithful...
This paper focuses on task-agnostic prompt compression for better generalizability and efficiency. Considering the redundancy in natural language, existing approaches compress prompts by removing...
https://arxiv.org/abs/2403.12968

LLMLingua-2 | Learn Compression Target via Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression
LLMLingua-2: Data distillation for efficient and faithful task-agnostic prompt compression (ACL 2024).
https://llmlingua.com/llmlingua2.html


Seonglae Cho