Last Layer Key-Value Cache is All You Need
COLM 2026 ER Workshop (Spotlight) · PDF
Keep one history for every layer. A standard Transformer stores a separate key–value history at each layer. LastKV builds a shared history from the final layer of completed chunks, letting every future layer reuse deeply processed context. This reduces repeated storage and feeds deep representations of the past back into future computation.
Within a chunk, tokens still use ordinary causal attention in parallel. At the boundary, the final-layer KV becomes shared history. A learned blend with the previous chunk’s layer-specific KV smooths the transition. For T tokens processed in complete chunks (T ≥ C), L layers and chunk size C, retained history contains T + (L−1)C KV positions instead of LT, approaching 1/L of the standard history as the sequence grows. These counts cover persistent KV storage.
Under matched long-context pretraining, LastKV has the highest average RULER score among the compared methods at 124M and 760M; the 124M score rises from 16.52 to 20.05 over the standard Transformer. Compared with a Transformer continued-pretrained on the same corpus for the same number of steps, the 8B LastKV model raises the mean across 17 selected benchmarks from 46.21 to 48.62, including HumanEval 36.46 to 50.30. The paper also evaluates state tracking, novel-view synthesis and teacher-forced video generation. Training overhead depends on chunk size.
