Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

Published in Preprint, 2026

We present Long-Horizon-Terminal-Bench (LHTB), a benchmark of 46 complex tasks across nine domains — experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing, among others. Unlike prior terminal-agent benchmarks, tasks are decomposed into graded subtasks with dense, reward-based grading, enabling partial credit and finer-grained evaluation. Tasks typically demand hundreds of episodes and minutes to hours of execution time, emphasizing sustained planning and iterative refinement. Across 15 state-of-the-art models, the best agent reaches only 15.2% success at relaxed thresholds and 10.9% at perfect accuracy, revealing substantial headroom for long-horizon autonomous agents.

* Equal contribution.

arXivCodeBlog

Recommended citation: Zongxia Li*, Zhongzhi Li*, Yucheng Shi*, Ruhan Wang, Junyao Yang, Zhichao Liu, Xiyang Wu, Anhao Li, Yue Yu, Ninghao Liu, Lichao Sun, Haotao Mi, Leowei Liang. (2026). "Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading." Preprint.
Download Paper