●builderYou can build more reliable agents by training them to write the very tool schemas they must eventually invoke.
●researcherThe use of three independent reward axes for schema, code, and outcome provides a more granular gradient for agent training.