LSC-DPO Improves Preference Optimization via Learning-Signal Control
October 7, 2026
LSC-DPO addresses the diminishing sensitivity of the standard DPO logistic loss as preference margins scale. By dynamically regulating the learning signal based on a geometric analysis of the sigmoid factor, the method improves performance on AlpacaEval 2, MT-Bench, and Anthropic-HH.
HOW THIS AFFECTS YOU
●
researcherYou can apply this technique to stabilize alignment training when dealing with large preference margins.