G2MAF Improves Multi-Agent Flow Policies via Test-Time Gradient Guidance
September 28, 2026
G2MAF uses a single globally normalized, projected critic gradient to refine joint actions for frozen multi-agent policies at test-time. The framework achieves mean relative gains of 9.2% on MPE and 8.9% on SMAC benchmarks by ensuring feasible and coordinated agent corrections.
HOW THIS AFFECTS YOU
●
builderThis offers a way to improve multi-agent coordination during inference without full policy retraining.
●
researcherYou can optimize frozen policies without retraining using test-time gradient refinement.