Skip to content

Author

Yunpeng Ba

We have 4 of 6 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

AIM-ZO: Activation-Informed Subspace Maintenance for Zeroth-Order LLM Fine-Tuning

Zeroth-order (ZO) optimization offers a memory-efficient alternative for LLM fine-tuning by estimating updates only from forward evaluations of perturbed parameters, without backpropagation or activation storage. However, in billion-parameter LLMs, isotropic perturbations often waste many forward evaluations on weakly...

Yue Xie, Zhi Zheng, Yun-Peng Ba et al. · 0 citations
Preprint Aug 2026

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging

This work proposes Hyper-ES, a subspace-based ES framework that avoids the weakness of ES in full-parameter search while exploiting its strength in low-dimensional optimization, and consistently outperforms GRPO-LoRA while requiring 10% fewer space-consuming gradient updates.

Yuntian Gu, Zhi Zheng, Yun-Peng Ba et al. · 0 citations
#machine learning Preprint Aug 2026

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

These findings position ES as a distinct reasoning post-training paradigm rather than a less effective, memory-efficient alternative to GRPO, and study how hyperparameter design affects the effectiveness of ES, demonstrating that ES requires a smaller population size in a larger LLM.

Yunpeng Ba, Zhi Zheng, Yue Xie et al. · 0 citations
#machine learning Preprint Aug 2026

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to fine-tune larger LLMs...

Zhi Zheng, Rong-Sheng Chen, Yunpeng Ba et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.