OPDSearch+: On-Policy Distillation for Search-Augmented Reasoning
August 26, 2026
OPDSearch+ enables search-augmented reasoning in small language models using a frozen, off-the-shelf instruct model as a teacher. This approach avoids the high cost of task-specific teacher fine-tuning while using RL to refine student policy distributions for better retrieval-based reasoning.
HOW THIS AFFECTS YOU
●
builderThis provides a pathway to deploy more capable, small-scale RAG reasoning models.
●
researcherYou can explore on-policy distillation without the overhead of training specialized teacher models.