Self-Play Search Distillation for Superhuman Reasoning Data
September 28, 2026
SPSD generates superhuman synthetic data by converting MuZero-like self-play search records from board games into structured chains-of-thought. Applying this to Qwen3-4B-Base increased mean mathematical benchmark scores from 24.1 to 36 through environment-grounded supervision.
HOW THIS AFFECTS YOU
●
builderYou can use environment-grounded self-play to generate high-quality reasoning data for specialized domains.
●
researcherThis demonstrates the transferability of search-based reasoning data from games to abstract mathematics.