BenchDrift Quantifies Performance Sensitivity to Problem Phrasing
August 13, 2026
The BenchDrift framework demonstrates that rephrasing benchmark problems across linguistic and structural axes causes significant correctness flips in model scores. Findings indicate that phrasing sensitivity does not diminish with model scale; instead, stronger models experience more significant performance losses from rephrasing.
HOW THIS AFFECTS YOU
●
researcherYou should account for phrasing-induced drift when evaluating model progress on benchmarks like MMLU or GSM8K.