An omni-modal model series that scales RL compute using asynchronous training with up to 3.7B tokens per step. It utilizes a hybrid-SWA architecture and diverse environments across code, visual, and cyber domains for self-improvement.
HOW THIS AFFECTS YOU
●
researcherYou can study how massive-scale RL compute impacts multimodal reasoning capabilities.
●
investorThis demonstrates the growing importance of RL compute scaling in the foundation model race.