MEND Optimizes Flow Models via Proximal Velocity Matching
October 5, 2026
MEND is a reinforcement learning method for flow models that uses proximal velocity matching to cap rewards within prompt groups. It outperforms Flow-GRPO in significantly fewer updates—100 vs 4,000—without requiring KL penalties or frozen reference models.
HOW THIS AFFECTS YOU
●
builderThis offers a much faster and less computationally expensive path for optimizing flow-based generative models.
●
researcherYou can achieve superior post-training of flow models using a more efficient, gradient-based velocity approach.