FlowBalance Calibrates Self-Improvement Using Verifier-Grounded Advantage
September 2, 2026
FlowBalance improves reasoning models by learning a normalized distribution over complete responses. It uses a frozen training-time view to calculate token-level log-probability gains, calibrating these scores with verifier-derived group advantage to prevent reinforcing false confidence on negative-advantage trajectories.
HOW THIS AFFECTS YOU
●
builderThis offers a more robust way to fine-tune reasoning models using their own outputs.
●
researcherYou can use this method to stabilize the inner loop of on-policy self-improvement.