Skip to content

Author

Peichao Lai

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

Cross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement Learning

Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credit. Ideally, step-level credit should incorporate evidence be...

Yang-Yang Ren, Hao-Dong Zhu, Lin-Lin Yang et al. · 0 citations
Jul 2026

DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models

DI performs Direct Preference Optimization (DPO) after supervised fine-tuning to strengthen task alignment with human preferences, and introduces a controlled decoding process that enforces fixed output formats and restricts predictions to candidate sets.

Yilei Wang, Jiaxin Gan, Kexuan Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.