模型为什么越走越偏:从 LLM 强化学习到自回归视频的 Drift 发布于: 2026 年 9 月 26 日从 Score Centering、BPO、MaxRL 到 Self Forcing、视频偏好优化与 KV 量化校正:用分布、反馈循环和干预位置,理解不同 domain 的 drift。