Skip to content
Preprint

Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents

Aug 2026 · 0 citations · 51 references
Computer Science

TL;DR

Experiments on GUI agent benchmarks show that LACL-GUI provides more effective learning signals and consistently improves agent performance over prior methods, highlighting the value of trajectory-level supervision in contrastive RLVR.

Abstract

Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) have shown strong potential for automating tasks across diverse digital environments, where reinforcement learning (RL) has become a dominant training paradigm. However, widely used methods such as Group Relative Policy Optimization (GRPO) suffer from reward-gradient misalignment, leading to inefficient and unstable optimization. Recent work addresses this issue by reformulating RL with verifiable rewards (RLVR) as contrastive or classification-based objectives, which improve stability by eliminating problematic gradient behaviors. Despite this progress, existing contrastive RLVR methods rely primarily on outcome-level supervision and fail to capture fine-grained differences in trajectory quality within the same outcome category. In this paper, we propose Length-Aware Contrastive Learning for GUI Agents (LACL-GUI), a contrastive RLVR framework that incorporates trajectory-level quality signals into policy optimization. LACL-GUI introduces structured preferences within both successful and failed trajectories, encouraging concise successful executions and differentiating failure quality based on divergence from successful trajectories, while preserving optimization stability. Experiments on GUI agent benchmarks show that LACL-GUI provides more effective learning signals and consistently improves agent performance over prior methods, highlighting the value of trajectory-level supervision in contrastive RLVR.

View source

Similar papers

#machine learning Preprint Aug 2026

Learning Generalizable Behaviors for Terminal Agents

River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization is proposed, which achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks.

Yi-Fan Yao, Bo Pang, Xuan-Phi Nguyen et al. · 1 citation
#artificial intelligence Preprint Aug 2026

WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement Learning

WM-R1 is the first reinforcement learning framework that trains mobile GUI agents with world models instead of real environments, eliminating the need for real-environment interaction, supports massively parallelized and step-level granularized trajectory generation grounded in world models, and introduces a multi-dime...

Yu Han, Tianwen Qian · 0 citations
Preprint Sep 2026

Learning from Reliable Negatives: Confidence-Anchored Test-Time Adaptation for GUI Grounding

Graphical User Interface (GUI) grounding is essential for autonomous agents to map natural language instructions to precise screen coordinates. However, existing supervised fine-tuning and reinforcement learning methods are constrained by the high cost of annotation, creating a scalability bottleneck. In this paper, we...

Yi-Zhou Liu, Fei Tang, Yuchen Yan et al. · 0 citations
#machine learning Preprint Sep 2026

Towards Scalable RLVR: Multimodal Instruction Following Data Synthesis and Distillation

Multimodal instruction following (MMIF) is crucial for building generalist agents. However, current training paradigms rely heavily on Supervised Fine-Tuning (SFT), which often leads to surface-level pattern matching and degrades general capabilities. While Reinforcement Learning with Verifiable Rewards (RLVR) offers a...

Yi-Rong Zeng, Sai Zhang, Yu-Xian Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

UnifiedPlayers, a cooperative framework comprising a Planning Player that generates tasks, an Execution Player that produces multi-turn trajectories with Python tool calls, and an Evaluation Player that constructs executable verifiers, highlights cooperation among specialized players as a promising path toward self-enh...

Wen-Jie Liao, Liang Zhao, Ze-Hong Cao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.