Mr.LHDR Benchmark for Multimodal Long-Horizon Research Agents
September 11, 2026
Mr.LHDR evaluates deep research agents on long, interdependent reasoning chains requiring an average of 12.1 intermediate conclusions. The benchmark integrates multimodal evidence, including video frames, maps, and charts, to test agentic capabilities in solving complex, multi-step research tasks.
HOW THIS AFFECTS YOU
●
builderYou can use this to stress-test the reasoning depth and multimodal integration of your autonomous research agents.
●
researcherThis provides a more rigorous metric for long-horizon dependency and reasoning depth than current benchmarks.