LongNovel Benchmark for Hallucination Detection in Long-Context Summarization
August 20, 2026
LongNovel provides a multi-scale bilingual benchmark for detecting hallucinations in long-context novel summarization. The dataset includes 29 Chinese novels ranging from 16k to 100k tokens and utilizes Multi-Model Arbitration to validate eight distinct hallucination types.
HOW THIS AFFECTS YOU
●
builderThis offers a specialized testing framework for long-context RAG and summarization applications.
●
researcherYou can use this to study how hallucination patterns evolve as context windows expand.