Biography

I am a Ph.D. student in Computer Science at the University of Wisconsin–Madison, advised by Junjie Hu.

I work on how a long-context model should hold on to what it has read. Attention keeps everything and pays for it; recurrent and fast-weight models compress and forget. Almost everything I do sits on that trade — what the memory is, how it is written, and how it is read back. I consider long context the most important open problem in LLMs.

My work runs in three lines:

Browse the Research directory for individual project pages with methods, results and paper figures.

Education

  • Ph.D., University of Wisconsin–Madison  ·  Sept. 2024 – June 2028 (expected)
  • M.S., Peking University  ·  Sept. 2022 – June 2024
  • B.S., Beijing Jiaotong University  ·  Sept. 2018 – June 2022

Experience

Awards

  • COLM 2025 Spotlight — PyramidKV

News

  • July 2026   MENTOR accepted to Findings of ACL 2026!
  • Oct. 2025   Gave a talk at COLM 2025 in Montreal on PyramidKV!
  • Sept. 2025   R-KV accepted by NeurIPS 2025!
  • Sept. 2025   COMMA accepted by TMLR!
  • Sept. 2025   PyramidKV selected as Spotlight at COLM 2025!
  • July 2025   PyramidKV accepted by COLM 2025!
  • June 2025   I joined Adobe Research as a research intern!
  • May 2025   AdaptiveStep accepted by ICML 2025!
  • May 2025   Math-Minos accepted by ACL 2025!
  • Jan. 2025   Four papers (DnD-Transformer, HeadKV, Omni-MATH, GDPO) accepted by ICLR 2025!
  • Aug. 2024   VeCAF accepted by ACM Multimedia 2024!
  • May 2024   Four papers (PCA-Bench, ZeroED, CENSOR, FairEval) accepted by ACL 2024!
  • Mar. 2024   Two papers (DialogVCS, ALSACE) accepted by NAACL 2024!
  • Jan. 2024   MMICL accepted by ICLR 2024!
  • May 2023   SANTA accepted by ACL 2023!
  • Jan. 2023   CTR accepted by ICLR 2023!

Selected Work

Full write-ups on the Research page · complete list on Publications.

Test-Time Training

Universal Test-Time Training
Shared fast-weight memory across time and depth, for language modeling and novel-view synthesis. At matched state and active compute, uTTT-MoE gains 2.6 and 2.1 RULER points over layer-private TTT-MoE at 124M and 760M.
PDF · Page · Code · Details

Test-Time Generation
Generate each memory read from fresh noise, guided by a query. TTG improves average retrieval and novel-view reconstruction over the tested fast-weight baselines, and repeated reads recover alternative values in a controlled memory test.
PDF · Details

Efficient Architecture

Last Layer Key-Value Cache is All You Need
Let every layer read a shared last-layer KV history, feeding deep past representations back into future computation. Chunked processing preserves parallel attention within each chunk and reduces long-term KV storage toward 1/L of a standard L-layer Transformer.
PDF · Details

KV Cache Compression

PyramidKV  COLM 2025 Spotlight
Zefan Cai, Yichi Zhang, Bofei Gao, Yuliang Liu, Yucheng Li, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Junjie Hu, Wen Xiao
Attention funnels with depth, so the cache budget should too. Full-cache quality at 12% of the KV cache.
PDF · Page · Code · Details

R-KV  NeurIPS 2025
Zefan Cai, Wen Xiao, Hanshi Sun, Cheng Luo, Yikai Zhang, Ke Wan, Yucheng Li, Yeyang Zhou, Li-Wen Chang, Jiuxiang Gu, Zhen Dong, Animashree Anandkumar, Abedelkadir Asi, Junjie Hu
Reasoning traces are self-repetitive, so evict for redundancy as well as importance. ~100% of full-cache accuracy at a 10% budget, 6.6× throughput.
PDF · Page · Code · Details