Skip to content

Author

Shuyue Hu

We have 7 of 25 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching

Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In t...

Yang Chen, Yi-Tan Zhang, Michael J. Witbrock et al. · 0 citations
#artificial intelligence Preprint Sep 2026

RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents

This paper introduces RSI-router, a routing framework that constructs subtask-level model assignments and model-specific skills through recursive self-improvement over accumulated experience and establishes a stronger performance--cost Pareto frontier than 9 routing methods.

Hao Li, Hang-Fan Zhang, Zhi-Yao Cui et al. · 0 citations
Preprint Aug 2026

SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

Experiments across diverse models, benchmarks, and agent harnesses show that supervised fine-tuning on SKT-generated trajectories consistently improves skill-use performance, establishing verified data synthesis as an effective and scalable approach for skill-use training.

Zelin Tan, Yi-Qun Zhang, Hao Li et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

In a multi-day deployment with more than 70 iterations, HoH autonomously develops a first-person-shooter game, featuring a coherent storyline, fully implemented core mechanics, human-playable experience, polished visuals and integrated audio.

Hao Yan, Min-Le Su, Hang-Fan Zhang et al. · 2 citations
Preprint Aug 2026

JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills

This work introduces \textsc{JailbreakSkill}, a skill-centric framework for scaling automated red-teaming through reusable and continuously evolving attack capabilities, which packages existing attack strategies into modular, agent-ready skills that can be directly reused and adaptively selected across tasks and target...

Xiaoyu Wen, Jiajia Li, Zhida He et al. · 2 citations
Preprint Aug 2026

AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions

This work presents AgentPanel, a multi-agent forum for human--AI collaboration in scientific exploration, a multi-agent forum for human--AI collaboration in scientific exploration that outperforms a centralized multi-agent debate baseline and shows that users value AgentPanel for perspective diversity and exploration s...

Zhi-Yao Cui, Qianyi Wang, Hao Yan et al. · 1 citation

Stabilizing Rubric Integration Training via Decoupled Advantage Normalization

We propose Process-Aware Policy Optimization (PAPO), a method that integrates process-level evaluation into Group Relative Policy Optimization (GRPO) through decoupled advantage normalization, to address two limitations of existing reward designs. Outcome reward models (ORM) evaluate only final-answer correctness, trea...

Zelin Tan, Zhouliang Yu, Bo-Cheng Lin et al. · 4 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.