●builderYou can reduce training costs and complexity by simulating environment responses directly within the policy training loop.
●researcherThis method offers a new way to optimize agentic policies through end-to-end joint optimization of actions and environment modeling.