Open-Weight Models Fail to Introspect Internal Computations
August 24, 2026
An evaluation of eight open-weight models using the Open-Weight Masked Introspection (OWMI) framework shows that models cannot accurately report on internal state changes. Across 78,000 measurements, models failed to distinguish between real interventions and sham perturbations, performing at chance levels.
HOW THIS AFFECTS YOU
●
researcherThis challenges the assumption that frontier models possess reliable mechanistic interpretability or self-auditing capabilities.