Skip to content

Author

Yunpeng Zhang

We have 2 of 2 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jun 2026

Data-Efficient Online Training for Direct Alignment in LLMs

In recent years, online Direct Alignment from Preferences (DAP) has emerged as a popular alternative for Reinforcement Learning from Human Feedback (RLHF) due to its training stability and simplicity. In online DAP, training relies on preference data, each composed of a question and a pair of large language model (LLM)...

Chi Zhang, Jia-Chen T. Wang, Kun He et al. · 0 citations
Preprint Jul 2026

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

This work progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduces LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift.

Fan Jiang, Zhaoxu Sun, Mengchao Wang et al. · 6 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.