Capacity-Dependent Data Selection for Reasoning SFT
August 17, 2026
Likelihood-based response selection for reasoning SFT follows a Fast-Fit/Slow-Gain pattern where high-likelihood data accelerates early improvements but impacts long-term gains differently based on model size. Experiments on 1.5B to 8B parameter models show that the effectiveness of teacher-generated supervision is strictly tied to student capacity and training duration.
HOW THIS AFFECTS YOU
●
builderYou may achieve faster convergence with high-likelihood data, but monitor long-term performance on smaller models.
●
researcherConsider model capacity as a primary variable when designing likelihood-based data selection pipelines.