PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

← All research

KV Cache Compression & Long-Context Inference

COLM 2025 · Spotlight · PDF · Page · Code

Spend the cache budget where information is still spread out. PyramidKV observes that attention changes across depth: lower layers read broadly, while higher layers concentrate on fewer tokens. It assigns larger KV caches to lower layers and smaller caches to upper layers, keeping the total budget fixed. Within each head, attention from a recent instruction window selects which tokens to retain. The method requires no model training.

Full, fixed-position, uniform-budget and PyramidKV caches, showing larger caches in lower layers.
Figure 1: PyramidKV redistributes a fixed cache budget across layers to follow attention's increasing concentration with depth. Click to enlarge.

The paper evaluates 17 LongBench datasets and Needle-in-a-Haystack retrieval with Llama-3-8B/70B-Instruct and Mistral-7B-Instruct. At a 2,048-token cache budget, Llama-3-8B-Instruct reaches a mean LongBench score of 41.49, compared with 41.46 for the full cache. Under a much smaller 64-token budget, its TREC accuracy is 58.00, versus 38.50 for SnapKV. These results show that allocating memory across layers can preserve quality under compression and help at tight budgets; the quality retained depends on the model, task and cache size.