Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning

← All research

KV Cache Compression & Long-Context Inference

ICLR 2025 (Poster) · PDF · Code · Evaluation Data · BABILong Data

Allocate memory to the heads that use it. HeadKV replaces fixed per-layer KV budgets with a shared budget distributed across individual attention heads. HeadKV-R2 estimates importance using retrieval questions that require a contextual reasoning step, then measures how much a head attends to the correct answer. At prefill, each head receives a basic allocation plus an importance-weighted share, and attention from the trailing instruction window selects the entries to retain. The model’s weights are unchanged.

HeadKV scores attention heads with a contextual reasoning probe and allocates a shared KV budget.
Figure 1: head scores reflect retrieval and contextual reasoning, and determine each head's share of the KV budget. Click to enlarge.

On six LongBench QA datasets with Llama-3-8B-Instruct, a 128-entry average cache budget yields 32.00 versus 32.90 for the full cache: about 97% of performance while retaining 1.5% of the average 8,683-token input. At the same budget, Ada-SnapKV scores 28.52. Experiments also cover Mistral-7B-Instruct, LooGLE, needle retrieval and contextual reasoning. The strongest gains appear at low cache budgets.