Skip to content

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

Sep 2026 · 0 citations · 44 references
Computer Science

TL;DR

This work introduces ActObs, which also supervises the observation tokens already present in each trajectory, and traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model.

Abstract

Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets. We ask whether this convention provides the best initialization for subsequent reinforcement learning. We introduce ActObs, which also supervises the observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy to model action consequences without adding data, parameters, sequence tokens, or forward passes. The methods perform similarly after SFT but diverge after GRPO. On Qwen3-4B, GRPO from ActObs achieves higher pass@k at every evaluated sampling budget than its action-only counterpart on Terminal-Bench 2.0. On Qwen3-8B, it trades some pass@1 reliability for higher pass@k (+3.4 pp at pass@16) and solves more distinct tasks. The advantage extends to cross-domain code editing on aider-polyglot (+4.2 pp at pass@1 at 4B), whose tasks are unseen during SFT and RL. ActObs retains more entropy during RL while requiring less policy movement, leaving the final policy closer to its SFT initialization. Our analysis traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model. Joint supervision prevents this one-sided specialization, preserving consequence prediction and preparing the policy for downstream exploration.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Does Learning to Predict the World Help Agents Act? Auditing World-Model Post-Training

Predicting how an environment will change before acting is a natural route to better decision making for agents. Recent post-training methods therefore require agents to predict the next observation and turn that prediction into a reward or a direct supervision signal, which is called world model. Existing next-observa...

Xin-Yu Che, Hang Yan, Yan-Chen Liu et al. · 0 citations
#machine learning Preprint Aug 2026

Three Steps at a Time: Learning Representations from Action Sequences in Contrastive RL

This work extends contrastive reinforcement learning (CRL), a prototypical self-supervised method, to operate over action chunks, and finds that this results in large, pervasive gains across established offline and online benchmarks: +31.7% and +93.1% across 18 and 11 environments respectively.

Michal Korniak, Kamil Dybek, Benjamin Eysenbach et al. · 1 citation · ⚡1
Preprint Sep 2026

Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models

FIND is introduced, an agentic real-world RL framework that closes the loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace and reframes autonomous practice as a scene-conditioned, performance-aware task-selection problem.

Yuan Fang, Ze-Chu Li, Hao-Lei Tong et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy d...

Yi-Tong Qiao, Tian-Tian He, Lei Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning

Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajecto...

Bo-Wen Zhang, Jun-Wei He, Mao-Qi Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier by executing each task's own verifier.

Jun-Yao Yang, Yu-Cheng Shi, Zhong-Zhi Li et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.