Skip to content

Author

Yangyang Ren

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

GraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM Agents

Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an action in sampled rollouts, but do not explicitly capture its r...

Hao-Dong Zhu, Yang-Yang Ren, Chang-Bai Li et al. · 0 citations
#machine learning Preprint Sep 2026

Cross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement Learning

Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credit. Ideally, step-level credit should incorporate evidence be...

Yang-Yang Ren, Hao-Dong Zhu, Lin-Lin Yang et al. · 0 citations
Jul 2026

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning

A Kalman-Guided Prompt Selection method (KGPS), which reformulates prompt selection as a dynamic state estimation problem rather than static difficulty prediction, and consistently improves both final accuracy and rollout efficiency over strong baselines, establishing state-of-the-art performance among online prompt se...

Hao-Dong Zhu, Yang-Yang Ren, Yanjing Li et al. · 2 citations · ⚡2

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.