●builderYou can improve search agent reliability by decoupling planning from synthesis to prevent context accumulation noise.
●researcherRDPO offers a new RL training method that combines high-level outcome rewards with granular, turn-level evaluations.