LitTraceQA Benchmark for Scientific Question Answering Verification
August 10, 2026
LitTraceQA evaluates multi-stage grounding by requiring systems to return canonical paper identifiers, specific evidence locations, and answers across formats like tables and equations. It targets the reliability of RAG systems in extracting evidence from complex scientific figures and text spans.
HOW THIS AFFECTS YOU
●
builderYou can use this benchmark to test the grounding accuracy of your scientific RAG pipelines.
●
researcherThis provides a structured way to evaluate how models handle non-textual evidence like equations and tables.