Incomplete Reference Sets Distort Calibration in Theory-of-Mind Tracking
August 27, 2026
Finite reference sets in Theory-of-Mind (ToM) evaluation create false negatives that reverse model calibration rankings. Using the NQ-open DPR-BERT pipeline, researchers found that exact-match metrics can improve reported calibration error by 0.045 while actually worsening human-aligned correctness by 0.074.
HOW THIS AFFECTS YOU
●
researcherYou should avoid using exact-match reference sets when evaluating open-ended reasoning or ToM capabilities.
●
policyBe cautious of benchmark results that claim high model alignment if they rely on limited reference sets.