Biography
I am a Ph.D. student in Computer Science at the University of Wisconsin–Madison, advised by Junjie Hu. My Chinese name is 蔡泽凡.
I work on how a long-context model should hold on to what it has read. Attention keeps everything and pays for it; recurrent and fast-weight models compress and forget. Almost everything I do sits on that trade — what the memory is, how it is written, and how it is read back. I consider long context the most important open problem in LLMs.
My work runs in three lines:
- Test-Time Training — memory as fast weights optimized online during the forward pass. Who owns them, what supervises them, and whether reading them should be deterministic at all.
- Efficient Architecture — removing redundant axes from memory: network depth in the KV cache, cumulative overlap in layer-to-layer routing.
- KV Cache Compression & Long-Context Inference — a fixed budget, and where to spend it: across layers, across heads, across a model’s own chain of thought.
Each project is written up in detail on the Research page — the bet it makes, what was built, where it was tested, and what came out.
Education
- Ph.D., University of Wisconsin–Madison · Sept. 2024 – June 2028 (expected)
- Advisor: Junjie Hu
- M.S., Peking University · Sept. 2022 – June 2024
- Advisor: Baobao Chang
- B.S., Beijing Jiaotong University · Sept. 2018 – June 2022
Experience
- Adobe Research — San Jose, California · May 2025 – July 2026
- Supervisor: Hao Tan
- Three projects on long-context memory: LastKV, Universal Test-Time Training, and Test-Time Generation.
- Microsoft Azure
- Supervisor: Wen Xiao
Awards
- COLM 2025 Spotlight (24 / 418) — PyramidKV
News
- July 2026 MENTOR accepted to Findings of ACL 2026!
- Oct. 2025 Gave a talk at COLM 2025 in Montreal on PyramidKV!
- Sept. 2025 R-KV accepted by NeurIPS 2025!
- Sept. 2025 COMMA accepted by TMLR!
- Sept. 2025 PyramidKV selected as Spotlight (24 / 418) at COLM 2025!
- July 2025 PyramidKV accepted by COLM 2025!
- June 2025 I joined Adobe Research as a research intern!
- May 2025 AdaptiveStep accepted by ICML 2025!
- May 2025 Math-Minos accepted by ACL 2025!
- Jan. 2025 Four papers (DnD-Transformer, HeadKV, Omni-MATH, GDPO) accepted by ICLR 2025!
- Aug. 2024 VeCAF accepted by ACM Multimedia 2024!
- May 2024 Four papers (PCA-Bench, ZeroED, CENSOR, FairEval) accepted by ACL 2024!
- Mar. 2024 Two papers (DialogVCS, ALSACE) accepted by NAACL 2024!
- Jan. 2024 MMICL accepted by ICLR 2024!
- May 2023 SANTA accepted by ACL 2023!
- Jan. 2023 CTR accepted by ICLR 2023!
Selected Work
Full write-ups on the Research page · complete list on Publications.
Test-Time Training
Universal Test-Time Training Adobe Research · under submission
Fast weights shared across the whole depth stack instead of one private state per layer. RULER retrieval rises from 8.5 to 15.5 at 124M and 23.2 to 27.8 at 760M across the routing-and-ownership ladder, of which the controlled sharing step contributes +2.6 and +2.1.
Details
Test-Time Generation Adobe Research · under submission
Memory readout as key-conditioned transport from noise rather than deterministic regression, so multimodal bindings stop collapsing to their mean. Scene-level rendering 16.27 → 17.46 dB.
Details
Test-Time Training with Next-Token Prediction Preprint 2026
Xuan Ouyang*, Zefan Cai*, Junjie Hu
Supervise the fast-weight write with the model’s own next-token signal instead of a learned local proxy. The only method improving RULER retrieval over base on all four backbones tested.
PDF · Details
Efficient Architecture
LastKV: Last Layer Key-Value Cache is All You Need Adobe Research · under submission
Commit only the last layer’s KV at each chunk boundary and let every future layer read it. An 8B Transformer converted to retain one layer of KV out of 32, at −0.6 average across 12 benchmarks.
Details
Delta Attention Residuals Preprint 2026
Cheng Luo*, Zefan Cai*, Junjie Hu
Layer-to-layer routing collapses because its sources are cumulative and near-identical. Routing over per-sublayer deltas instead: 8.2% perplexity improvement at 7.57B, where the prior method is worse than doing nothing.
PDF · Details
KV Cache Compression
PyramidKV COLM 2025 Spotlight (24 / 418)
Zefan Cai, Yichi Zhang, Bofei Gao, Yuliang Liu, Yucheng Li, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Junjie Hu, Wen Xiao
Attention funnels with depth, so the cache budget should too. Full-cache quality at 12% of the KV cache.
PDF · Page · Code · Details
R-KV NeurIPS 2025
Zefan Cai, Wen Xiao, Hanshi Sun, Cheng Luo, Yikai Zhang, Ke Wan, Yucheng Li, Yeyang Zhou, Li-Wen Chang, Jiuxiang Gu, Zhen Dong, Anima Anandkumar, Abedelkadir Asi, Junjie Hu
Reasoning traces are self-repetitive, so evict for redundancy as well as importance. ~100% of full-cache accuracy at a 10% budget, 6.6× throughput.
PDF · Page · Code · Details
