Active Taskless Distillation Transfers Capabilities via Single-Word Teacher Signals
September 25, 2026
Active Taskless Distillation (ATD) enables capability transfer between models using only a single word per prompt from the teacher. In experiments with Qwen2.5-1.5B, this method achieved a 5.34 percentage point gain on HumanEval+ without requiring target-task examples or teacher logits.
HOW THIS AFFECTS YOU
●
builderYou may be able to distill capabilities into smaller models using much more efficient, single-token signals.
●
researcherThis provides a new method for probing behavioral shadows and performing extremely low-bandwidth distillation.