Salt++: Context-Aligned Post-Training for Streaming Multimodal Generation
September 28, 2026
Salt++ is a two-stage post-training framework designed to improve few-step streaming audio-video generation. It uses Causal Self-Flow and context-aligned autoregressive Distribution Matching Distillation to resolve mismatches between generation and scoring contexts in causal models.
HOW THIS AFFECTS YOU
●
builderYou can implement more stable and efficient streaming multimodal models using this distillation recipe.
●
researcherThis framework addresses specific context-related challenges in causal modeling for audio-video tasks.