●builderThe scale of RL compute and multi-modal training in this series sets a new baseline for agent performance in complex environments.
●researcherThis scaling approach provides a blueprint for using agentic grading and high-throughput asynchronous training to drive model self-improvement.