Universal Test-Time Training

Zefan Cai1,*   Qinzhe Hu1,*   Ziqiao Ma2   Hao Tan3   Junjie Hu1

1University of Wisconsin–Madison   2University of Michigan   3Adobe    *Equal contribution

Code

TL;DR — Test-Time Training compresses context into fast weights, and existing designs make those fast weights layer-private: every layer owns a state no other layer touches. We argue this is an arbitrary restriction and study universal fast weights, where the whole depth stack shares one state or one pool of states — uTTT-Dense, a shared dense operator, and uTTT-MoE, a globally shared expert pool, with TTT-MoE, a routed pool that stays layer-local, as the controlled midpoint.

Sharing also changes what the model is: one state persisting across depth makes the layer index a second recurrence axis, so an L-layer stack unrolls into a single recurrence over (chunk, layer) pairs, and state written deep at one chunk is read shallow at the next. Layer-private TTT is the fixed-partition special case of that model.

Across 32K-context language modeling and LVSM-style novel view synthesis, moving fast weights from layer-private, to layer-local routed, to a globally shared pool improves long-context retrieval monotonically — RULER accuracy from 8.5 to 15.5 at 124M and 23.2 to 27.8 at 760M, and object-level rendering from 24.5 to 25.5 dB — though the advantage concentrates at shorter contexts and does not yet hold at the longest evaluated length.

The ownership ladder

Universal Test-Time Training method overview
Universal Test-Time Training. Left: the model is a stack of N uTTT blocks. Middle: a block keeps its own normalization, projections, and local window attention private, and sends its projected queries, keys, and values to the uTTT memory path. Right: the memory itself is shared across the whole stack — uTTT-MoE (top) exposes it as a pool of fast-weight experts that every block addresses through a shared top-K router, and uTTT-Dense (bottom) exposes it as a single dense fast-weight operator that every block reads and updates in full. Blue arrows are reads, red arrows are updates.

Two task domains

Novel view synthesis  uttt_nvs/

  • LVSM-style rendering on Objaverse/GSO and DL3DV, both at 256×256
  • PSNR / SSIM / LPIPS full-test evaluation with scene bootstrap
  • uTTT-Dense (r1/r2/r4/r8) and uTTT-MoE, plus TTT-Dense and TTT-MoE
  • The ladder is a configuration switch on one implementation — 64 configs

Language modeling  uttt_llm/

  • 32K context, 124M and 760M, global batch 32
  • Per-token loss on Book-3 and RULER retrieval evaluation
  • uTTT-MoE plus LaCT, TTT-Dense, TTT-MoE and attention baselines — 57 configs
  • One model package per paper row, shipped as run; uTTT-Dense not yet released

Key results

Language modeling

Per-token loss versus RULER retrieval at 124M and 760M
Per-token loss versus long-context retrieval at (a) 124M and (b) 760M. Each point is one model; the upper-left corner is best, and the arrows trace the ownership ladder. Making the fast-weight pool universal (uTTT-MoE) moves the model toward that corner relative to its balanced and unbalanced layer-local counterparts (TTT-MoE w/ and w/o LB) and to the two layer-private dense baselines, the published LaCT and our own TTT-Dense, at both scales. The full-attention Transformer reaches higher retrieval at the cost of a key–value cache that grows with context.

Long-context language modeling at 32K. PTL is mean next-token loss at positions ≥30K on Book-3; RULER is accuracy averaged over four evaluation lengths. Bold marks the best bounded-state model; the full-attention Transformer† is a reference whose key–value cache grows with context.

ModelPTL ↓ (124M)RULER ↑ (124M)PTL ↓ (760M)RULER ↑ (760M)
Transformer† (full attention)3.086916.522.597834.16
LaCT3.12348.502.578723.21
TTT-Dense3.121912.652.580926.05
TTT-MoE w/o LB3.10759.992.584324.86
TTT-MoE w/ LB3.106212.822.581525.73
uTTT-MoE3.085715.462.576827.85
Per-token validation loss by position at 124M and 760M
Per-token validation loss by position within a 32K-token sequence, at (a) 124M and (b) 760M; PTL is the mean of the right-hand end (positions ≥ 30K). Loss discriminates the models far less than retrieval does, and more so with scale: at 760M the five fast-weight models fall within a 0.0075-nat band while their RULER spread is 4.6 points, and the two orderings no longer agree.

Sharing the bank is also the fastest of the three regimes. One global pool is written by a single grouped kernel per chunk, where the layer-local pools pay for L small updates; measured at 124M with matched activation checkpointing and full backward through the recurrent state:

RegimeExperts storedtok/s/GPU ↑min / 1K steps ↓PTL ↓RULER ↑
TTT-Dense (layer-private)129,61436.93.121912.65
TTT-MoE w/ LB (layer-local)423,88845.73.106212.82
uTTT-MoE (universal)4832,78233.33.085715.46

The largest pool is also the fastest. Throughput is implementation-sensitive: three implementations of the same 48-expert global pool span 30,644–32,782 tok/s/GPU, so cross-architecture throughput comparisons should be discounted accordingly.

Novel view synthesis

PSNR, SSIM and LPIPS against view index on GSO and DL3DV
Ownership against view index, full-test evaluation. TTT-Dense keeps one private dense state per layer, TTT-MoE-e8a1 a layer-local pool of 8 experts, and uTTT-MoE-e64a1 a global pool of 64 experts — the last two hold identical total capacity and parameter count. Rows are PSNR, SSIM, and LPIPS; left column Objaverse/GSO, right column DL3DV.

Novel view synthesis, full-test PSNR (mean over target views 1–23, then over scenes; 1,019 GSO scenes and 140 DL3DV scenes).

ModelPSNR ↑ (Objaverse/GSO)PSNR ↑ (DL3DV)
TTT-Dense23.9916.14
TTT-MoE-e8a1 (layer-local)24.0316.05
uTTT-MoE-e64a1 (global, top-1)24.6816.12
uTTT-MoE-e64a8 (global, top-8)25.5616.66
Transformer† (full attention)25.3617.41

Numbers from the paper's current version; full tables, confidence intervals, and the balancing and capacity sweeps are in the paper and reproducible from the released configs.

Novel view synthesis demos

uTTT-MoE renderings on held-out scenes, teacher-forced conditioning on ground-truth views 0…k−1.

Objaverse/GSO
object renderings · uTTT-MoE
coming soon
DL3DV
scene renderings · uTTT-MoE
coming soon