PARCEL Benchmark for Detecting Legal Hallucinations
October 9, 2026
The PARCEL benchmark evaluates LLM ability to verify legal claims against source texts using 3,396 labeled examples. While top models reach 0.97 accuracy, they consistently struggle to detect unsupported claims and fabricated but plausible citations.
HOW THIS AFFECTS YOU
●
builderYou should implement secondary verification layers for legal-domain applications.
●
researcherThe results highlight a specific failure mode in natural language inference regarding missing support.