●builderThis approach could significantly improve the training efficiency of agents performing multi-step tool-calling tasks.
●researcherThis provides a way to densify reward signals in agentic environments without relying on potentially biased learned reward models.