Recursive Synthesis for Long-Horizon Terminal Tasks
Published in Preprint, 2026
High-quality long-horizon training data for terminal agents is expensive to produce because each task must keep the instruction, environment, reference solution, and verifier mutually consistent. We present Recursive Synthetic Terminal Tasks (RST), a recursive verified synthesis framework that starts from verified seed tasks, extends the reference solution, realigns the verifier and instruction, validates the result in a fresh sandbox, and reuses accepted tasks as seeds. Across fifteen recursive rounds, RST produces 37,484 synthesized terminal-agent tasks at roughly $0.05 per task, with difficulty climbing as DeepSeek-V4-Pro pass@4 drops from 90% at R1 to 2.5% at R15. Fine-tuning on rejection-sampled trajectories from these tasks improves Qwen3.5 models by up to 10 points on Terminal-Bench 2, Terminal-Bench Hard, and Long-Horizon Terminal Bench.
* Equal contribution.
| arXiv | Project Page | Hugging Face |
Recommended citation: Zhongzhi Li*, Yucheng Shi*, Zongxia Li*, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang. (2026). "Recursive Synthesis for Long-Horizon Terminal Tasks." Preprint.
Download Paper
