Concurrent actions in large language model (LLM) agent environments require arbitration even when each proposal is individually valid. We implement a typed snapshot-settlement contract and audit three distinct properties: order sensitivity, useful progress, and replay consistency. Five settlement policies are tested in...
Hao-Tian Chen, Bo-Wen Ye, Yu-Ning Zhang et al.· 0 citations
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task re...
Jin-Hao Dong, Liang Zhao, Zi-Hao Yue et al.· 0 citations
Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better...
Bo-Wen Ye, Lei Li, Shi-Cheng Li et al.· 0 citations
The first fully automated framework that synthesizes high-quality, proof-centric benchmarks from natural language mathematical corpora and a new type of hybrid-formatted questions, named ``$m$-out-of-$n$ multiple judge questions'', specifically designed to enable robust, automatic evaluation while being resilient to gu...
Ye-Bo Peng, Zixiang Liu, Yao-Ming Li et al.· arXiv.org· 1 citation
The core of SSVAL is Visual Anchor Prompt Injection (VAPI), which introduces prompts that absorb rich knowledge from external VFMs during training, enabling them to serve as stable visual anchors that mitigate representation deviation during inference.
Qian-Long Yang, Bowen Ye, Xianda Guo et al.· 0 citations
CoEvo-Mem alternates between updating the router with the memory bank fixed and evolving the memory bank with the retrieval policy fixed, demonstrating the importance of retrieval-memory coevolution.
Bowen Ye, Yongchao Xu, Zhijian Li et al.· 0 citations
This work introduces PersonaForge, a user simulation framework for synthesizing realistic multi-turn user--agent interactions that combines a four-dimensional persona space, SOUL-driven behavioral control calibrated to real-user statistics, and Reverse Deep Construction grounded in authentic seed queries.
Hanglong Lv, Dawei Zhu, Lei Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.