Inspect Evals census reveals claim-replay failures
August 24, 2026
A formal analysis of 124 Inspect Evals units found that 110 units cannot be deterministically replayed due to missing historical evidence or semantic grounding. The study formalizes a claim-replay layer using a frozen substrate and grounded families to identify discrepancies between reported metrics and verifiable computation.
HOW THIS AFFECTS YOU
●
researcherYou should account for missing semantic grounding when evaluating the reproducibility of LLM benchmarks.
●
policyThis highlights the need for standardized licensing and data provenance in model evaluation artifacts.