Multi-Head Attention Residuals
Preprint 2026 · co-first author · PDF · Website · Code · MHAR 8B checkpoint · Control 8B checkpoint
Let different feature subspaces read different layers. Attention residuals give each sublayer a learned mixture of earlier layer outputs, but a single routing distribution makes every feature channel read the same layers in the same proportions. Multi-Head Attention Residuals (MHAR) splits the routing query into heads, allowing each feature subspace to choose its own mixture of the depth history. One head recovers ordinary attention residuals; splitting the existing query adds no parameters.
In the paper’s matched-step experiments on a quality-filtered Nemotron-based corpus, validation loss falls by 0.061, 0.149 and 0.140 relative to a standard Transformer at 100M, 350M and 1B parameters. The best head counts are four or eight in the tested settings. An 8B conversion using the delta residual form starts with a closed routing gate to preserve the pretrained model, then learns routing during continued training. Against a control with the same schedule, data order and token budget, it improves GSM8K by 3.2 percentage points and GPQA by 3.1 points.
Reading the depth history still costs computation. Fused Triton kernels raise end-to-end training throughput to 0.55–0.88× the standard Transformer in the reported 100M–1B settings, with approximately baseline peak memory. The plain-Transformer comparison uses equal training steps; the paper leaves a direct comparison at equal wall-clock time to future work.
