PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
KV Cache Compression & Long-Context Inference
COLM 2025 · Spotlight · PDF · Page · Code
Spend the cache budget where information is still spread out. PyramidKV observes that attention changes across depth: lower layers read broadly, while higher layers concentrate on fewer tokens. It assigns larger KV caches to lower layers and smaller caches to upper layers, keeping the total budget fixed. Within each head, attention from a recent instruction window selects which tokens to retain. The method requires no model training.
The paper evaluates 17 LongBench datasets and Needle-in-a-Haystack retrieval with Llama-3-8B/70B-Instruct and Mistral-7B-Instruct. At a 2,048-token cache budget, Llama-3-8B-Instruct reaches a mean LongBench score of 41.49, compared with 41.46 for the full cache. Under a much smaller 64-token budget, its TREC accuracy is 58.00, versus 38.50 for SnapKV. These results show that allocating memory across layers can preserve quality under compression and help at tight budgets; the quality retained depends on the model, task and cache size.
