Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable answers. Yet the policy and reward system form a dynamic feedback loop: as the policy optimizes the current reward, an initially useful reward system may become unreliable due to reward hacking or reduced respon...
Weiyuan Li, Aili Chen, Xin-Tao Wang et al.· 0 citations
This work proposes SocialRL, a multi-turn reinforcement learning framework using PPO that propagates delayed outcome rewards back to each turn, enabling long-horizon planning and demonstrates the effectiveness of SocialRL across synthetic and real social scenes, as well as standard and challenging social scenarios.
Jia-Ning Wang, Xin-Tao Wang, Ai-Li Chen et al.· 1 citation
Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than accuracy-based metrics, because predictions that are numerically close to the ground truth can still violate operational constraints, combine fields in physically inconsistent ways, or fail to produce usable str...
Ya-Lun Wu, Jun-Feng Fang, Jia-Wei Wang et al.· 0 citations
This work presents FORGE (Fabricated Orchestrated Reasoning chain for aGent Exploitation), a two-level attack that combines intra-document reasoning fabrication with inter-document chain coordination to hijack subtask planning.
Yu-Ru Pan, Ziheng Zhang, Junxiang Lei et al.· arXiv.org· 1 citation
Results show that EMAS can turn experience from new samples into reusable updates to MAS topology and prompts, and is best or tied in six of eight model--benchmark settings.
Chao Fei, Qing-Yi Si, Kai-Huan Liang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.