Internal Probes Reveal Knowledge Hidden by Saturated Decision Thresholds
September 7, 2026
A 0.6B parameter model demonstrates behavioral failure in logical verification despite internal probes achieving 0.96 AUC on correct verdicts. Research shows that a single scalar threshold offset of +4.6 sigma suppresses correct knowledge from the model's output logits.
HOW THIS AFFECTS YOU
●
researcherYou should investigate logit-based decision thresholds rather than relying solely on behavioral accuracy to assess model knowledge.