F2DR introduces a fine-grained reward framework that evaluates DeepSearch workflows across three dimensions: Content, Trajectory, and Answer. Accompanied by the RM-Bench benchmark, it addresses the failure of static single-turn reward models to capture the complexities of iterative planning and retrieval loops.
HOW THIS AFFECTS YOU
●
builderYou can use this framework to train more reliable reward models for agentic search pipelines.
●
researcherThis provides a more granular evaluation metric for iterative, multi-step reasoning processes.