Command-Conditioned PPO for Real-Time Strategy Agents
October 9, 2026
This method separates strategic command selection from unit control in MicroRTS using a constrained PPO executor and a Thompson-sampling bandit strategist. This decoupling allows agents to adapt to diverse opponents by switching strategies without retraining the execution policy.
HOW THIS AFFECTS YOU
●
researcherYou can improve agent robustness by decoupling high-level strategic planning from low-level execution.