Pyramid-JiT architecture achieves faster text-to-image training without a VAE
October 6, 2026
Pyramid-JiT uses a decoder-only pixel space architecture to predict target images at multiple resolutions along a DiT trunk. This method matches Linum v2 FD-DINOv2 scores using 11.3x fewer training samples and 4.3x fewer GPU-hours while processing 4x the pixels.
HOW THIS AFFECTS YOU
●
builderYou can reduce training costs and inference latency by bypassing the VAE bottleneck.
●
researcherThis demonstrates scaling decoder-only pixel space architectures via multi-resolution target prediction.