LLaDA-Image: 6B Diffusion Transformer with Image-Only Pre-training
September 2, 2026
LLaDA-Image pairs a 6B Diffusion Transformer with a frozen LLaDA2.0-Mini language backbone. It utilizes image-only pre-training on 220M samples and the Muon optimizer to achieve high photorealism and fine-grained instruction following.
HOW THIS AFFECTS YOU
●
builderYou can use a unified architecture that combines diffusion language models with DiTs for better instruction following.
●
researcherThe use of image-only mid-training and the Muon optimizer offers a new recipe for scalable DiT training.