Skip to content

Author

Changwei Wang

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning

Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's...

Zeng-Huang Fu, Zhao-Yang Li, Qiu-Yuan Ai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SIPO: Selective-Inference Policy Optimization for Tree-Structured Agentic RL

Tree-structured reinforcement learning trains search agents by comparing alternative continuations and propagating terminal rewards to intermediate decisions. Adaptive expansion, however, creates a statistical asymmetry: an incumbent is selected using its own generation statistic, whereas fresh siblings are sampled aft...

Zeng-Huang Fu, Ning Chen, Ming-Da Jia et al. · 0 citations
Jul 2026

UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following

The Unified Embodied Seeking and Following Benchmark (UESF-Bench) is introduced, a large-scale and diverse benchmark for embodied human seeking and following that requires agents to handle semantic-guided exploration, reliable behavior switching and recovery, and delayed identity grounding.

Kun Yu, Jianhua Yang, Yixiang Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.