This 191.43B-token synthetic corpus uses a dual generation strategy to optimize small model pre-training in STEM domains. It employs a weak student model to generate corrective explanations from failures and contrastive reasoning from successes across 19 different domains.
HOW THIS AFFECTS YOU
●
builderYou can use this high-density data to improve the STEM reasoning of edge-scale models.
●
researcherThe targeted teacher distillation method offers a new way to curate synthetic training signals.