LLMLingua
LLMLingua compresses prompts using budget allocation and iterative token-level compression. It selects information to retain instead of simply deleting text because it is old.
Source: LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models.
Referenced in An Empirical Study of Harness Design for Coding Agents.