T1: 122B MoE Model for Long-Horizon Terminal Tasks
September 11, 2026
T1 is a 122B Mixture-of-Experts model trained via reinforcement learning to execute long-horizon tasks in cloud sandboxes for over 300 tool-call turns. The training uses a dense process reward system and TITO construction to stabilize optimization during agentic rollouts.
HOW THIS AFFECTS YOU
●
builderThis provides a recipe for training agents capable of sustaining much longer, complex tool-use sequences.
●
researcherThe use of TITO construction and rollout routing replay offers new methods for stabilizing MoE training in agentic settings.