T1: 122B Mixture-of-Experts Model for Long-Horizon Tool Use
September 9, 2026
T1 is a 122B parameter MoE model trained via reinforcement learning for tasks requiring over 300 tool-call turns. It uses TITO construction and rollout routing replay to stabilize training in real shell cloud sandboxes.
HOW THIS AFFECTS YOU
●
builderYou can build agents capable of sustained, complex tool-use sequences in real environments.
●
researcherThe TITO construction and drift repair methods offer a recipe for stable long-horizon RL training.