Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning
KV Cache Compression & Long-Context Inference
ICLR 2025 (Poster) · PDF · Code · Evaluation Data · BABILong Data
Allocate memory to the heads that use it. HeadKV replaces fixed per-layer KV budgets with a shared budget distributed across individual attention heads. HeadKV-R2 estimates importance using retrieval questions that require a contextual reasoning step, then measures how much a head attends to the correct answer. At prefill, each head receives a basic allocation plus an importance-weighted share, and attention from the trailing instruction window selects the entries to retain. The model’s weights are unchanged.
On six LongBench QA datasets with Llama-3-8B-Instruct, a 128-entry average cache budget yields 32.00 versus 32.90 for the full cache: about 97% of performance while retaining 1.5% of the average 8,683-token input. At the same budget, Ada-SnapKV scores 28.52. Experiments also cover Mistral-7B-Instruct, LooGLE, needle retrieval and contextual reasoning. The strongest gains appear at low cache budgets.
