Qwen3 Models Reach 98.7% Accuracy in Robust Numerical Verification
October 2, 2026
Adversarial fine-tuning on numerically perturbed examples allows small Qwen3 models (0.6B–8B) to achieve 98.7% accuracy on label-flipping tasks. This outperforms frontier systems like GPT-5.4 Pro and Gemini 2.5 Flash, which scored approximately 74% on the same metric.
HOW THIS AFFECTS YOU
●
builderYou can use parameter-efficient fine-tuning to make small models significantly more reliable for math-heavy applications.
●
researcherYou can apply this fine-tuning recipe to improve robustness against numerical decision boundary shifts.