DR-to-Long: Boosting Long-Context Understanding in Research Agents
September 21, 2026
The DR-to-Long method improves deep research agents by repurposing RL trajectories into long-context QA training data. By replacing compact summaries with full web content, the approach targets the 61.6% of prediction errors caused by long-context hallucinations and evidence integration failures.
HOW THIS AFFECTS YOU
●
builderYou can improve your agent's ability to synthesize information from multiple long documents using this training strategy.
●
researcherThis provides a new way to bridge the data gap for long-context training using existing agent trajectories.