[arXiv]score: 0.17
Extremely Sparse Supervision Incentivizes Reasoning Ability
September 7, 2026
On-policy distillation using as few as one or two tokens per reasoning trajectory can match or surpass full-token training performance. Testing on the Qwen3 family shows that optimizing over 0.05% of generated tokens effectively incentivizes reasoning capabilities, suggesting that dense supervision at every token is unnecessary for effective post-training.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy