●builderYou can improve the efficiency of agent training by implementing backtracked clue recovery to reward useful intermediate steps.
●researcherThis method addresses the limitations of uniform SFT/RL in long-horizon tasks by providing granular reward signals.