Small 0.8B Models Gain 3.8x QA Performance via RLVR
September 25, 2026
Training Qwen3.5-0.8B using Group Relative Policy Optimization (GRPO) with an interleaved Wikipedia-search tool achieves a 0.352 exact match on MuSiQue. This demonstrates that verifiable rewards can improve open-domain question answering in sub-1B parameter models without requiring distillation from larger teachers.
HOW THIS AFFECTS YOU
●
builderYou can use RLVR and search-based rewards to significantly boost the reasoning capabilities of small, edge-deployable models.
●
researcherThis extends the applicability of the reason-over-search recipe to the sub-billion parameter regime.