Universal Test-Time Training
COLM 2026 ER Workshop (Spotlight) · PDF · Page · Code
One memory, shared across depth. Most test-time training models give each layer its own fast weights: a compact memory updated as the model reads. uTTT makes that memory available to every layer while keeping each layer’s backbone parameters separate. Deep layers can write information that shallow layers read in later chunks.
The paper develops two versions: uTTT-MoE, which routes reads and writes to a shared expert pool, and uTTT-Dense, which uses a shared dense operator. All layers read the state from the start of a chunk; their writes are combined at the boundary. This preserves parallel processing within the chunk.
At matched total fast-weight state and active compute, uTTT-MoE improves average RULER retrieval over layer-private TTT-MoE from 12.82 to 15.46 at 124M and 25.73 to 27.85 at 760M, and trains up to 1.39× faster in the measured configurations. In object-level novel-view synthesis, routed sharing adds 0.92 dB at the 23rd view. Dense sharing adds 0.76 dB while using one-eighth of the fast-weight state at matched per-layer compute. Average language retrieval remains below full attention; the gains over private memory vary with context length.
