PROOF Benchmark Profiles Factuality Gaps in Open-Weight Models
September 25, 2026
The PROOF benchmark uses 18,486 Wikidata-grounded questions to profile factual coverage across 101 classes and 392 properties. Evaluation of 18 open-weight models shows base factual accuracy between 6.58% and 57.59%, with domain-specific performance spreads reaching 36.4 percentage points.
HOW THIS AFFECTS YOU
●
researcherYou can use this to identify specific semantic domains where your model's factual retrieval fails.