Skip to content

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

Sep 2026 · 1 citation · 32 references
Computer Science

TL;DR

T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier by executing each task's own verifier.

Abstract

Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier. We provide a comprehensive recipe: First, an aggressively warm-started to stabilize actor-critic training, with a dense process reward scoring trajectories by the absolute number of passing verifiers. Second, stable optimization through TITO construction, training on the exact sampled token identifiers with drift repair at turn boundaries, and rollout routing replay, recording the sampler's per-token expert choices at every MoE layer and replaying them during training. Third, fully out-of-distribution training corpus: isolated seeds and synthesized tasks disjoint from Terminal-Bench 2.1 ensures gains reflect genuine capability transfer over benchmark overfitting. Together, TITO and R3 cut the training-to-inference log-probability difference from 0.021 to 0.013, with exactly aligned zero token drift in the loss region. On Terminal-Bench 2.1, our post-train pipeline raises initial base model from 43.8% to T1 with 64.0% resolved. On Long-Horizon Terminal Bench, T1 reaches 27.9% and surpasses GPT-5.4 and GLM-5.1.

View source

Similar papers

Preprint Aug 2026

Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks

SINKFLEX-RL, a modular training system for RL in dual-control tool-use environments that combines a Gymnasium-compatible environment wrapper, a VERL-style rollout dataflow, group-relative policy optimization without a separate value model, and a sink-aware FlexAttention path designed to preserve model-specific sink sca...

Zelei Cheng, Amritansh Mishra, Sambit Sahu et al. · 0 citations
#natural language process... Preprint Sep 2026

Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents

Long-horizon coding agents receive verifiable rewards only after completing expensive sequences of tool calls. This increases inference cost, amplifies early wrong hypotheses, and can lead to sparse terminal reward and unstable training. We introduce Contextual Early Reward (CER), which predicts terminal reward through...

Ji-Han Yao, Si-Han Zeng, Shang-Bin Feng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

CANOPY (Coverage-ANchored On-PolicY RL), a minimalist protocol attacking both Signal starvation and policy drift, internalizes long-horizon capability directly into a small open model; the complete training stack is planned to be released at https://github.com/AlibabaResearch/SignalCoverageRL.

Li-Ming Pu, Xiao-Xiao Li, Yi-Fu Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization

Designing dense reward functions for low-level reinforcement learning (RL) control remains difficult. Recent work uses large language models (LLMs) to iteratively generate and refine reward functions using policy-training feedback within scripted search algorithms. However, evaluating each candidate requires a full RL...

Ming-Hao Li, Rui Tan, Rui-Hang Wang · 0 citations
Preprint Aug 2026

Harness-RL: Black-Box Reinforcement Learning with Action-Args Decoupling for Central-Agent Multi-Agent Harnesses

Harness-RL is introduced, a structured reinforcement learning framework that combines Conflict-Aware Policy Optimization (CAPO) with interface-level black-box trajectory construction and supports both central-only and joint multi-agent training.

Xin-Ke Jiang, Zhi-Xin Zhang, Zhi-Bang Yang et al. · 1 citation
Preprint Sep 2026

SeeQ: Training Generalist Value Functions for Long-Horizon Robotic Manipulation

Despite rapid progress, generalist robot policies remain brittle on complex, long-horizon tasks that comprise multiple stages or require repeated attempts and deliberation on the same underlying stage before success. Q-value functions can improve these policies by ranking candidate actions or guiding policy improvement...

Saksham Singh, Zhe-Yuan Hu, Max Sobol Mark et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.