Delta Attention Residuals
Preprint, arXiv 2026 · co-first author · PDF · Code
Route what each layer adds. Cumulative hidden states share much of their content, making them difficult to distinguish when attention routes across depth. Delta Attention Residuals instead use the changes introduced by attention and MLP sublayers as routing sources. Their weighted combination is added to the current residual stream. The per-sublayer variant keeps individual contributions; Delta Block groups them to reduce the number of sources and the training cost.
In the reported 10,000-step FineWeb-Edu pretraining experiments, delta routing improves validation perplexity across tested architectures from 220M to 7.57B parameters. At 7.57B, Delta Block reaches 16.00 perplexity, compared with 17.43 for standard residuals and 18.58 for the paper’s Attention Residuals reimplementation: an 8.2% reduction relative to standard residuals. This configuration uses 589.8K extra routing parameters, with 35% lower training throughput and 3% more peak GPU memory than the baseline. On the Qwen3-0.6B routing diagnostic, average maximum attention weight increases from 0.35 to 0.62, supporting more selective cross-layer routing.
