Biography
I am a Ph.D. student in Computer Science at the University of Wisconsin–Madison, advised by Junjie Hu.
I work on how a long-context model should hold on to what it has read. Attention keeps everything and pays for it; recurrent and fast-weight models compress and forget. Almost everything I do sits on that trade — what the memory is, how it is written, and how it is read back. I consider long context the most important open problem in LLMs.
My work runs in three lines:
Browse the Research directory for individual project pages with methods, results and paper figures.
Education
- Ph.D., University of Wisconsin–Madison · Sept. 2024 – June 2028 (expected)
- Advisor: Junjie Hu
- M.S., Peking University · Sept. 2022 – June 2024
- Advisor: Baobao Chang
- B.S., Beijing Jiaotong University · Sept. 2018 – June 2022
Experience
- Adobe Research — San Jose, California · May 2025 – July 2026
- Supervisor: Hao Tan
- Three projects on long-context memory: LastKV, Universal Test-Time Training, and Test-Time Generation.
- Microsoft Azure
- Supervisor: Wen Xiao
Awards
- COLM 2025 Spotlight — PyramidKV
News
- July 2026 MENTOR accepted to Findings of ACL 2026!
- Oct. 2025 Gave a talk at COLM 2025 in Montreal on PyramidKV!
- Sept. 2025 R-KV accepted by NeurIPS 2025!
- Sept. 2025 COMMA accepted by TMLR!
- Sept. 2025 PyramidKV selected as Spotlight at COLM 2025!
- July 2025 PyramidKV accepted by COLM 2025!
- June 2025 I joined Adobe Research as a research intern!
- May 2025 AdaptiveStep accepted by ICML 2025!
- May 2025 Math-Minos accepted by ACL 2025!
- Jan. 2025 Four papers (DnD-Transformer, HeadKV, Omni-MATH, GDPO) accepted by ICLR 2025!
- Aug. 2024 VeCAF accepted by ACM Multimedia 2024!
- May 2024 Four papers (PCA-Bench, ZeroED, CENSOR, FairEval) accepted by ACL 2024!
- Mar. 2024 Two papers (DialogVCS, ALSACE) accepted by NAACL 2024!
- Jan. 2024 MMICL accepted by ICLR 2024!
- May 2023 SANTA accepted by ACL 2023!
- Jan. 2023 CTR accepted by ICLR 2023!
Selected Work
Full write-ups on the Research page · complete list on Publications.
Test-Time Training
Universal Test-Time Training
Shared fast-weight memory across time and depth, for language modeling and novel-view synthesis. At matched state and active compute, uTTT-MoE gains 2.6 and 2.1 RULER points over layer-private TTT-MoE at 124M and 760M.
PDF · Page · Code · Details
Test-Time Generation
Generate each memory read from fresh noise, guided by a query. TTG improves average retrieval and novel-view reconstruction over the tested fast-weight baselines, and repeated reads recover alternative values in a controlled memory test.
PDF · Details
Efficient Architecture
Last Layer Key-Value Cache is All You Need
Let every layer read a shared last-layer KV history, feeding deep past representations back into future computation. Chunked processing preserves parallel attention within each chunk and reduces long-term KV storage toward 1/L of a standard L-layer Transformer.
PDF · Details
KV Cache Compression
PyramidKV COLM 2025 Spotlight
Zefan Cai, Yichi Zhang, Bofei Gao, Yuliang Liu, Yucheng Li, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Junjie Hu, Wen Xiao
Attention funnels with depth, so the cache budget should too. Full-cache quality at 12% of the KV cache.
PDF · Page · Code · Details
R-KV NeurIPS 2025
Zefan Cai, Wen Xiao, Hanshi Sun, Cheng Luo, Yikai Zhang, Ke Wan, Yucheng Li, Yeyang Zhou, Li-Wen Chang, Jiuxiang Gu, Zhen Dong, Animashree Anandkumar, Abedelkadir Asi, Junjie Hu
Reasoning traces are self-repetitive, so evict for redundancy as well as importance. ~100% of full-cache accuracy at a 10% budget, 6.6× throughput.
PDF · Page · Code · Details
