[HUGGINGFACE]score: 0.42
Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
October 6, 2026
Iris-3B is a 3B-parameter pixel-space text-to-image transformer pretrained using a 256 to 1024 resolution curriculum. While designed to avoid VAE compression losses, empirical testing shows the model provides no significant improvement over latent-space generative priors for monocular depth estimation or image restoration tasks.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy