MMLongBench-Doc-V2 Corrects Annotation Errors in Long-Document QA
August 5, 2026
MMLongBench-Doc-V2 updates the benchmark by correcting 106 annotations and replacing string-match metrics with a semantic LLM judge. The revision addresses issues where incorrect or ambiguous ground-truth labels previously penalized models for correct answers.
HOW THIS AFFECTS YOU
●
builderUse this updated dataset to more accurately test your RAG or long-context document processing systems.
●
researcherYou can now use a more reliable benchmark for evaluating long-context retrieval and reasoning.