Skip to content

Author

Yanghua Xiao

We have 8 of 23 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning

Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable answers. Yet the policy and reward system form a dynamic feedback loop: as the policy optimizes the current reward, an initially useful reward system may become unreliable due to reward hacking or reduced respon...

Weiyuan Li, Aili Chen, Xin-Tao Wang et al. · 0 citations
#natural language process... Preprint Sep 2026

SocialRL: Refining LLMs'Social Intelligence through Multi-turn Reinforcement Learning and Reward Design

This work proposes SocialRL, a multi-turn reinforcement learning framework using PPO that propagates delayed outcome rewards back to each turn, enabling long-horizon planning and demonstrates the effectiveness of SocialRL across synthetic and real social scenes, as well as standard and challenging social scenarios.

Jia-Ning Wang, Xin-Tao Wang, Ai-Li Chen et al. · 1 citation
#artificial intelligence Preprint Sep 2026

DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems

Large language model (LLM)-based multi-agent systems have experienced rapid growth in recent years. Despite their promise, such systems remain fragile, frequently exhibiting reasoning and coordination errors that can lead to system-level failures. Failure attribution in such systems relies on tracing natural language i...

Ze-Hao Wang, Lanjun Wang, Shi-Long Jin et al. · 1 citation
Conference Open access Jun 2026

Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents

This paper introduces FinED-Bench, the first publicly public benchmark for FinED-Bench, which covers nine real-world financial scenarios, and includes over 900 documents reported in 2025 that are unseen by existing language models.

Ying He, Zhouhong Gu, Zhecheng Hu et al. · 2 citations
Preprint Aug 2026

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

Skill-Use is introduced, a benchmark that evaluates skill use under progressive disclosure, where an agent sees only a skill's name and short description and must retrieve the full procedure before following it.

Jinyi Han, Yuanjian Xu, Ying Liao et al. · 3 citations · ⚡1
Preprint Jun 2026

From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training

A novel policy gradient method is introduced, Hindsight Policy Optimization (HPO), that projects both the current policy distribution and the hindsight distribution into an intent space and extracts low-variance learning signals from the Wasserstein distance between them.

Zishang Jiang, Tingyun Li, Jinyi Han et al. · 0 citations
Jul 2026

From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models

A multi-layer taxonomy of 14 capability domains and 91 subskills across Primitive, Constructed, and Integrative layers is introduced and supports research organization, coverage audits, evaluation interpretation, and testable hypotheses for diagnosis, training, and transfer.

Shixin Fang, Jiachen Wo, Wen-Juan Qin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.