OnlineQAT Achieves Superior Performance for 2-bit and 3-bit LLMs
October 8, 2026
OnlineQAT uses on-policy distillation to improve ultra-low-bit quantization by training on student-generated responses rather than fixed teacher outputs. On Qwen3-1.7B, the method achieves 57.28 at W3A16 and 32.52 at W2A16, outperforming ReasoningQAT in both configurations.
HOW THIS AFFECTS YOU
●
builderYou can deploy smaller, more efficient models with higher accuracy by using on-policy training signals.
●
researcherThe findings demonstrate that correcting quantization errors requires conditioning on student-generated prefixes.