LLaDA-Image: 6B Diffusion Transformer with Open Training Recipes
September 4, 2026
A unified framework pairing a 6B DiT with a frozen vision-language backbone trained via image-only pre-training. The model achieves 53.53 on Qwen-Image-Bench and includes a distilled Turbo version capable of 2-4 sampling steps.
HOW THIS AFFECTS YOU
●
builderYou can utilize the Turbo version for high-speed, high-fidelity image generation.
●
researcherYou can replicate or build upon the image-only pre-training pipeline.