Test-Time Generation

← All research

Test-Time Training

COLM 2026 ER Workshop (Spotlight) · PDF

Read memory by generating from it. A compact memory can associate several possible values with the same query. A deterministic read trained with squared loss may average those values. Test-Time Generation (TTG) instead learns from noisy key–value examples and generates each read from fresh noise, guided by the query and the stored fast weights. The memory stays fixed during a read and updates when new context arrives.

Conceptual comparison of deterministic conditional-mean readout and stochastic generative readout from a fixed-size fast-weight memory.
Conceptual illustration from the paper: a deterministic conditional mean can sit between plausible values; generative readout aims to represent alternatives. The curves illustrate the conceptual model. Click to enlarge.

Across the reported 124M, 760M and 3B model families, TTG achieves the highest mean RULER retrieval accuracy among the compared models with bounded memory state. On the full GSO object and DL3DV scene test sets, average reconstruction PSNR rises from LaCT’s 23.99 / 16.14 dB to 24.93 / 16.59 dB. A Transformer with a growing KV cache remains stronger on these reconstruction benchmarks.

A controlled test stores two values per key. Repeated stochastic reads recover both at low and moderate loads, showing that a stored state can support alternative outputs. At the highest tested load all methods fail, and adding readout steps does not improve the language model trained for one step.

Selected ladybug and cactus reconstructions: ground truth, deterministic reference, and TTG across multiple views.
Selected examples from the paper, with all 24 views written to memory and observed views included in evaluation. The displayed gains describe these two cases; the full-test averages are reported above. Click to enlarge.