Delta Attention Residuals

← All research

Efficient Architecture

Preprint, arXiv 2026 · co-first author · PDF · Code

Route what each layer adds. Cumulative hidden states share much of their content, making them difficult to distinguish when attention routes across depth. Delta Attention Residuals instead use the changes introduced by attention and MLP sublayers as routing sources. Their weighted combination is added to the current residual stream. The per-sublayer variant keeps individual contributions; Delta Block groups them to reduce the number of sources and the training cost.

Standard additive residuals, attention over cumulative states, and additive routing over sublayer deltas.
Figure 2: Delta Attention Residuals route over sublayer changes and add the result to the residual stream. Click to enlarge.

In the reported 10,000-step FineWeb-Edu pretraining experiments, delta routing improves validation perplexity across tested architectures from 220M to 7.57B parameters. At 7.57B, Delta Block reaches 16.00 perplexity, compared with 17.43 for standard residuals and 18.58 for the paper’s Attention Residuals reimplementation: an 8.2% reduction relative to standard residuals. This configuration uses 589.8K extra routing parameters, with 35% lower training throughput and 3% more peak GPU memory than the baseline. On the Qwen3-0.6B routing diagnostic, average maximum attention weight increases from 0.35 to 0.62, supporting more selective cross-layer routing.