FlowCPO provides offline preference alignment for flow and diffusion models
September 10, 2026
FlowCPO introduces an offline forward-KL objective that uses preferred and dispreferred samples without requiring online rollouts. It provides a tractable, non-negative surrogate loss for linear interpolation, addressing instability in previous methods like FlowDPO.
HOW THIS AFFECTS YOU
●
researcherYou can perform preference alignment on flow models using fixed datasets without the need for expensive online sampling.