StalePO Mitigates Model Degradation During Preference Optimization on Outdated Data
September 16, 2026
StalePO addresses the stale preference problem in machine translation where newer models are optimized using human post-edits from legacy systems. The method enforces downward likelihood movement on both responses and anchors the policy to its base response using token-level KL constraints to prevent quality erosion.
HOW THIS AFFECTS YOU
●
builderThis allows you to reuse expensive human preference data even after upgrading your base translation model.
●
researcherYou can leverage this objective to prevent DPO from overfitting to outdated human corrections.