ComPO Zeroth-Order Method for LLM Preference Alignment
September 15, 2026
Comparison-based Preference Optimization (ComPO) is a zeroth-order alignment method that uses comparison oracles to extract directional information from preference pairs. It circumvents the need for optimizing differentiable preference losses directly.
HOW THIS AFFECTS YOU
●
researcherYou can use this method to align models when preference data has small likelihood margins.