Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable answers. Yet the policy and reward system form a dynamic feedback loop: as the policy optimizes the current reward, an initially useful reward system may become unreliable due to reward hacking or reduced respon...
Weiyuan Li, Aili Chen, Xin-Tao Wang et al.· 0 citations
This work proposes SocialRL, a multi-turn reinforcement learning framework using PPO that propagates delayed outcome rewards back to each turn, enabling long-horizon planning and demonstrates the effectiveness of SocialRL across synthetic and real social scenes, as well as standard and challenging social scenarios.
Jia-Ning Wang, Xin-Tao Wang, Ai-Li Chen et al.· 1 citation
Large language model (LLM)-based multi-agent systems have experienced rapid growth in recent years. Despite their promise, such systems remain fragile, frequently exhibiting reasoning and coordination errors that can lead to system-level failures. Failure attribution in such systems relies on tracing natural language i...
Ze-Hao Wang, Lanjun Wang, Shi-Long Jin et al.· 1 citation
AnySearch is proposed, a framework that enables a single policy to perform budget-aware search under any budget constraint through a training scaffold and curriculum reinforcement learning, and outperforms baselines across all budget scales.
Xiaowei Sun, Jin Li, Yi-Li Hong et al.· 1 citation
This paper introduces FinED-Bench, the first publicly public benchmark for FinED-Bench, which covers nine real-world financial scenarios, and includes over 900 documents reported in 2025 that are unseen by existing language models.
Ying He, Zhouhong Gu, Zhecheng Hu et al.· Annual Meeting of the Associ...· 2 citations
Skill-Use is introduced, a benchmark that evaluates skill use under progressive disclosure, where an agent sees only a skill's name and short description and must retrieve the full procedure before following it.
A novel policy gradient method is introduced, Hindsight Policy Optimization (HPO), that projects both the current policy distribution and the hindsight distribution into an intent space and extracts low-variance learning signals from the Wasserstein distance between them.
Zishang Jiang, Tingyun Li, Jinyi Han et al.· 0 citations
A multi-layer taxonomy of 14 capability domains and 91 subskills across Primitive, Constructed, and Integrative layers is introduced and supports research organization, coverage audits, evaluation interpretation, and testable hypotheses for diagnosis, training, and transfer.