FlowBalance Calibrates Self-Improvement via Verifier-Grounded Trajectory Guidance
September 4, 2026
FlowBalance improves reasoning models by learning a normalized distribution over complete responses using a verifier-grounded self-improvement method. It uses a frozen training-time policy to produce token-level log-probability gains, calibrated by group advantage to prevent reinforcing false confidence or narrow solution modes.
HOW THIS AFFECTS YOU
●
researcherYou can use this method to stabilize the inner loop of on-policy reasoning reinforcement learning.