Scaling and Distilling Embeddings for Diffusion Language Models
October 2, 2026
Scaling embedding models from T5 to T5Gemma-2 improves diffusion language model performance, but highly discriminative embeddings can hinder continuous diffusion sampling. The authors propose distilling large teacher encoders into student encoders that better capture decoded probabilities to improve generation stability.
HOW THIS AFFECTS YOU
●
builderThis provides a path toward more stable diffusion-based text generation by optimizing the underlying embedding space.
●
researcherYou can apply this distillation method to create more diffusible latent spaces for continuous text generation.