Graphical User Interface (GUI) grounding is essential for autonomous agents to map natural language instructions to precise screen coordinates. However, existing supervised fine-tuning and reinforcement learning methods are constrained by the high cost of annotation, creating a scalability bottleneck. In this paper, we...
Yi-Zhou Liu, Fei Tang, Yuchen Yan et al.· 0 citations
A Gaussian guidance framework that draws the depth per task from a Gaussian whose center and spread are estimated online from rollouts already collected for policy optimization, requiring no probe rollouts or learned depth predictor is proposed.
Zi-Xuan Wang, Yan-Rui Miao, Zhengxi Lu et al.· 0 citations
Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories int...
Xing-Yu Wu, Yuchen Yan, Zhengxi Lu et al.· 0 citations
This work proposes RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD.
Yan Yu, Zhengxi Lu, Yi-Zhou Liu et al.· 1 citation
Web agents that act from rendered pixels avoid the fragility and heavy token cost of reading a page's HTML or accessibility tree, but training them depends on large amounts of high-quality interaction trajectories, and how to produce such data at scale remains an open problem. Public datasets typically contain only a f...
Fei Tang, Hua-Wen Shen, Zhiqiong Lu et al.· 0 citations
Automatic translation quality metrics trained on general-domain corpora systematically fail on social media content, where communicative intent is encoded in culturally loaded expressions (internet slang, homophonic ciphers, and platform-specific idioms) rather than surface token patterns. We conduct a systematic empir...
Yi-Wen Qiu, Linjuan Wu, Ding-Ming Li et al.· 0 citations
This work presents EasySteer, a unified framework for high-performance, extensible LLM steering built on vLLM, and demonstrates its effectiveness in overthinking mitigation, hallucination reduction, and other key applications.
Haolei Xu, Xinyu Mei, Yuchen Yan et al.· arXiv.org· 14 citations
Test-Time Policy Optimization is proposed, an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL and Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors.
Ao-Han Wang, Zhengxi Lu, Jianze Wang et al.· 0 citations
This work introduces PaperGym, a unified framework that turns each research paper into a complete training environment, and releases the pipeline, the 20,000-instance corpus PaperGym-20k, and the benchmarks PaperGym-Innov and PaperGym-Design.
Yu-Han Wang, Zhengxi Lu, Yuchen Yan et al.· 1 citation
This work proposes AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning that aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space.
Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao et al.· 7 citations· ⚡1
EMPO achieves substantial gains over the base model and achieves competitive performance against strong baselines such as UI-TARS-7B and GPT-4o, demonstrating better generalization than prior single-turn RL approaches.
Zhengxi Lu, Jiabo Ye, Fei Tang et al.· Annual Meeting of the Associ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.