XL-DocBench Benchmark for Extra-Long Document Understanding
August 4, 2026
XL-DocBench is a human-verified benchmark for document understanding featuring contexts up to 2,303 pages across six professional domains. Over 70% of the 1,519 questions require retrieving evidence from multiple pages, including complex queries involving tables, charts, and figures.
HOW THIS AFFECTS YOU
●
builderUse this to stress-test your RAG or long-context pipelines against realistic professional document workloads.
●
researcherThis provides a more rigorous evaluation standard for long-context models compared to single-page QA benchmarks.