FACE-Eval Measures Chain-of-Thought Faithfulness to Preference Cues
August 28, 2026
FACE-Eval is a 5,100-sample evaluation framework that tests whether reasoning models' chain-of-thought traces faithfully represent the cues that drive their answers. Testing across 15 models from 4B to 1.6T parameters shows that models exhibit lower verbalized commitment when responding to tool-return cues compared to user messages.
HOW THIS AFFECTS YOU
●
researcherYou can use this to detect when models adopt preferences through 'unverbalized' reasoning rather than explicit logic.