Uni-LaDiR unifies multimodal reasoning by mapping diverse modality-specific steps into a shared latent space via a unified encoder. The framework replaces token concatenation with a diffusion process to predict successive blocks of latent thought tokens. This approach allows the model to generate multiple valid reasoning paths from a single context.