Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In t...
Yang Chen, Yi-Tan Zhang, Michael J. Witbrock et al.· 0 citations
This paper introduces RSI-router, a routing framework that constructs subtask-level model assignments and model-specific skills through recursive self-improvement over accumulated experience and establishes a stronger performance--cost Pareto frontier than 9 routing methods.
Hao Li, Hang-Fan Zhang, Zhi-Yao Cui et al.· 0 citations
Experiments across diverse models, benchmarks, and agent harnesses show that supervised fine-tuning on SKT-generated trajectories consistently improves skill-use performance, establishing verified data synthesis as an effective and scalable approach for skill-use training.
Zelin Tan, Yi-Qun Zhang, Hao Li et al.· 2 citations
In a multi-day deployment with more than 70 iterations, HoH autonomously develops a first-person-shooter game, featuring a coherent storyline, fully implemented core mechanics, human-playable experience, polished visuals and integrated audio.
Hao Yan, Min-Le Su, Hang-Fan Zhang et al.· 2 citations
This work introduces \textsc{JailbreakSkill}, a skill-centric framework for scaling automated red-teaming through reusable and continuously evolving attack capabilities, which packages existing attack strategies into modular, agent-ready skills that can be directly reused and adaptively selected across tasks and target...
Xiaoyu Wen, Jiajia Li, Zhida He et al.· 2 citations
This work presents AgentPanel, a multi-agent forum for human--AI collaboration in scientific exploration, a multi-agent forum for human--AI collaboration in scientific exploration that outperforms a centralized multi-agent debate baseline and shows that users value AgentPanel for perspective diversity and exploration s...
Zhi-Yao Cui, Qianyi Wang, Hao Yan et al.· 1 citation
We propose Process-Aware Policy Optimization (PAPO), a method that integrates process-level evaluation into Group Relative Policy Optimization (GRPO) through decoupled advantage normalization, to address two limitations of existing reward designs. Outcome reward models (ORM) evaluate only final-answer correctness, trea...
Zelin Tan, Zhouliang Yu, Bo-Cheng Lin et al.· arXiv.org· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.