T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
Published in Preprint, 2026
We introduce T1, a 122B Mixture-of-Experts terminal agent trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task and rewarded by executing each task’s own verifier. The recipe combines an aggressively warm-started actor-critic with dense process rewards, TITO construction with drift repair at turn boundaries, and rollout routing replay for MoE training-inference consistency, using a fully out-of-distribution synthesized task corpus. On Terminal-Bench 2.1, the post-train pipeline raises the base model from 43.8% to 64.0% resolved. On Long-Horizon Terminal Bench, T1 reaches 27.9% and surpasses GPT-5.4 and GLM-5.1.
* Equal contribution.
| arXiv | Project Page | Hugging Face |
Recommended citation: Junyao Yang, Yucheng Shi, Zhongzhi Li, Ruhan Wang, Zongxia Li, Haitao Mi, Leowei Liang. (2026). "T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks." Preprint.
Download Paper
