Web agents that act from rendered pixels avoid the fragility and heavy token cost of reading a page's HTML or accessibility tree, but training them depends on large amounts of high-quality interaction trajectories, and how to produce such data at scale remains an open problem. Public datasets typically contain only a f...
Fei Tang, Hua-Wen Shen, Zhiqiong Lu et al.· 0 citations
This work presents EasySteer, a unified framework for high-performance, extensible LLM steering built on vLLM, and demonstrates its effectiveness in overthinking mitigation, hallucination reduction, and other key applications.
Haolei Xu, Xinyu Mei, Yuchen Yan et al.· arXiv.org· 14 citations
Test-Time Policy Optimization is proposed, an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL and Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors.
Ao-Han Wang, Zhengxi Lu, Jianze Wang et al.· 0 citations
Results show that self-contamination is a trainable component of the LiC gap, and propose MAIGO, an on-policy self-distillation method that reduces this contamination using history-cleaned references from the model's own policy.
This work introduces PaperGym, a unified framework that turns each research paper into a complete training environment, and releases the pipeline, the 20,000-instance corpus PaperGym-20k, and the benchmarks PaperGym-Innov and PaperGym-Design.
Yu-Han Wang, Zhengxi Lu, Yuchen Yan et al.· 1 citation
EMPO achieves substantial gains over the base model and achieves competitive performance against strong baselines such as UI-TARS-7B and GPT-4o, demonstrating better generalization than prior single-turn RL approaches.
Zhengxi Lu, Jiabo Ye, Fei Tang et al.· Annual Meeting of the Associ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.