Biography

I am a Ph.D. student in Computer Science at the University of Wisconsin–Madison, advised by Junjie Hu. My Chinese name is 蔡泽凡.

I work on how a long-context model should hold on to what it has read. Attention keeps everything and pays for it; recurrent and fast-weight models compress and forget. Almost everything I do sits on that trade — what the memory is, how it is written, and how it is read back. I consider long context the most important open problem in LLMs.

My work runs in three lines:

  • Test-Time Training — memory as fast weights optimized online during the forward pass. Who owns them, what supervises them, and whether reading them should be deterministic at all.
  • Efficient Architecture — removing redundant axes from memory: network depth in the KV cache, cumulative overlap in layer-to-layer routing.
  • KV Cache Compression & Long-Context Inference — a fixed budget, and where to spend it: across layers, across heads, across a model’s own chain of thought.

Each project is written up in detail on the Research page — the bet it makes, what was built, where it was tested, and what came out.

Education

  • Ph.D., University of Wisconsin–Madison  ·  Sept. 2024 – June 2028 (expected)
  • M.S., Peking University  ·  Sept. 2022 – June 2024
  • B.S., Beijing Jiaotong University  ·  Sept. 2018 – June 2022

Experience

Awards

  • COLM 2025 Spotlight (24 / 418) — PyramidKV

News

  • July 2026   MENTOR accepted to Findings of ACL 2026!
  • Oct. 2025   Gave a talk at COLM 2025 in Montreal on PyramidKV!
  • Sept. 2025   R-KV accepted by NeurIPS 2025!
  • Sept. 2025   COMMA accepted by TMLR!
  • Sept. 2025   PyramidKV selected as Spotlight (24 / 418) at COLM 2025!
  • July 2025   PyramidKV accepted by COLM 2025!
  • June 2025   I joined Adobe Research as a research intern!
  • May 2025   AdaptiveStep accepted by ICML 2025!
  • May 2025   Math-Minos accepted by ACL 2025!
  • Jan. 2025   Four papers (DnD-Transformer, HeadKV, Omni-MATH, GDPO) accepted by ICLR 2025!
  • Aug. 2024   VeCAF accepted by ACM Multimedia 2024!
  • May 2024   Four papers (PCA-Bench, ZeroED, CENSOR, FairEval) accepted by ACL 2024!
  • Mar. 2024   Two papers (DialogVCS, ALSACE) accepted by NAACL 2024!
  • Jan. 2024   MMICL accepted by ICLR 2024!
  • May 2023   SANTA accepted by ACL 2023!
  • Jan. 2023   CTR accepted by ICLR 2023!

Selected Work

Full write-ups on the Research page · complete list on Publications.

Test-Time Training

Universal Test-Time Training  Adobe Research · under submission
Fast weights shared across the whole depth stack instead of one private state per layer. RULER retrieval rises from 8.5 to 15.5 at 124M and 23.2 to 27.8 at 760M across the routing-and-ownership ladder, of which the controlled sharing step contributes +2.6 and +2.1.
Details

Test-Time Generation  Adobe Research · under submission
Memory readout as key-conditioned transport from noise rather than deterministic regression, so multimodal bindings stop collapsing to their mean. Scene-level rendering 16.27 → 17.46 dB.
Details

Test-Time Training with Next-Token Prediction  Preprint 2026
Xuan Ouyang*, Zefan Cai*, Junjie Hu
Supervise the fast-weight write with the model’s own next-token signal instead of a learned local proxy. The only method improving RULER retrieval over base on all four backbones tested.
PDF · Details

Efficient Architecture

LastKV: Last Layer Key-Value Cache is All You Need  Adobe Research · under submission
Commit only the last layer’s KV at each chunk boundary and let every future layer read it. An 8B Transformer converted to retain one layer of KV out of 32, at −0.6 average across 12 benchmarks.
Details

Delta Attention Residuals  Preprint 2026
Cheng Luo*, Zefan Cai*, Junjie Hu
Layer-to-layer routing collapses because its sources are cumulative and near-identical. Routing over per-sublayer deltas instead: 8.2% perplexity improvement at 7.57B, where the prior method is worse than doing nothing.
PDF · Details

KV Cache Compression

PyramidKV  COLM 2025 Spotlight (24 / 418)
Zefan Cai, Yichi Zhang, Bofei Gao, Yuliang Liu, Yucheng Li, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Junjie Hu, Wen Xiao
Attention funnels with depth, so the cache budget should too. Full-cache quality at 12% of the KV cache.
PDF · Page · Code · Details

R-KV  NeurIPS 2025
Zefan Cai, Wen Xiao, Hanshi Sun, Cheng Luo, Yikai Zhang, Ke Wan, Yucheng Li, Yeyang Zhou, Li-Wen Chang, Jiuxiang Gu, Zhen Dong, Anima Anandkumar, Abedelkadir Asi, Junjie Hu
Reasoning traces are self-repetitive, so evict for redundancy as well as importance. ~100% of full-cache accuracy at a 10% budget, 6.6× throughput.
PDF · Page · Code · Details