Large reasoning models (LRMs) have shown exceptional performance in complex tasks such as mathematics and coding. In the field of machine translation (MT), reinforcement learning (RL) has been utilized to enhance the quality of translations. However, traditional RL approaches rely heavily on the base model’s inherent...
Zengkui Sun, Jia-Li Zeng, Jiaan Wang et al.· Transactions of the Associat...· 0 citations
Long-horizon software engineering (SWE) agents trained with reinforcement learning with verifiable rewards (RLVR) typically receive only terminal outcome supervision, making it difficult to distinguish productive actions from redundant exploration or functional regressions. We propose SWE-MILE, an asynchronous potentia...
Chao-Qun Cui, Hao Zhou, Mei-Qi Chen et al.· 0 citations
Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction experience into persistent model capabilities. Existing self-improvement paradigms remain fragmented: test-time methods can explicitly extract e...
Shijie Ren, Xiting Wang, Meng Li et al.· 0 citations
This work presents VERA-RL, a reinforcement-learning formulation for scientific error detection over academic papers, and constructs VERA-13K, a 12,900-sample dataset organized into 4,300 matched chains, covering 6 scientific-error categories across the research workflow and broad natural-science domains.
Rongjin Li, Yuanxin Liu, Hao Zhou et al.· 0 citations
EvoBrowseComp is introduced, an evolving benchmark of 400 English and 400 Chinese contamination-free complex questions synthesized via live-web traversal that establishes a scalable paradigm for auto-updatable, high-difficulty benchmarking that keeps pace with both evolving world knowledge and advancing agent capabilit...
This work introduces PolyWorkBench, a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows, and proposes a hybrid framework that combines structural grading, executable verification, and LLM-based semantic assessment to enable comprehensive evaluation.
Benchmark evaluations reveal that agent performance varies substantially across languages and drops sharply on the harder cross-lingual tasks, and analysis shows that multilingual execution exposes systematic failure modes across planning, tool interaction, and decision-making in long-horizon agents.
Hongliang Li, Yijin Liu, Zhiwei Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.