Skip to content

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses

Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interactions grow, useful evidence can be buried by redundant or outdated context, making context management itself a key bottleneck. We introduce...

Wei-Yuan Li, Jing-Heng Xu, Ai-Li Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

OSPD: On-Policy Self-Distillation for Persona-Consistent Dialogue

Maintaining persona consistency across multi-turn dialogues remains a core challenge for role-playing language models. Off-policy distillation from external teachers incurs distribution mismatch that compounds across dialogue turns, while reinforcement learning struggles with reward ambiguity inherent in subjective per...

Rui Xu, Yikai Zhang, Ai-Li Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning

Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable answers. Yet the policy and reward system form a dynamic feedback loop: as the policy optimizes the current reward, an initially useful reward system may become unreliable due to reward hacking or reduced respon...

Weiyuan Li, Aili Chen, Xin-Tao Wang et al. · 0 citations
#natural language process... Preprint Sep 2026

SocialRL: Refining LLMs'Social Intelligence through Multi-turn Reinforcement Learning and Reward Design

This work proposes SocialRL, a multi-turn reinforcement learning framework using PPO that propagates delayed outcome rewards back to each turn, enabling long-horizon planning and demonstrates the effectiveness of SocialRL across synthetic and real social scenes, as well as standard and challenging social scenarios.

Jia-Ning Wang, Xin-Tao Wang, Ai-Li Chen et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.