Why Models Drift: From LLM Reinforcement Learning to Autoregressive Video
Published:
From Score Centering, BPO, and MaxRL to Self Forcing, video preference optimization, and KV quantization correction: understanding drift across domains through distributions, feedback loops, and points of intervention.
