Zeroth-order (ZO) optimization offers a memory-efficient alternative for LLM fine-tuning by estimating updates only from forward evaluations of perturbed parameters, without backpropagation or activation storage. However, in billion-parameter LLMs, isotropic perturbations often waste many forward evaluations on weakly...
Yue Xie, Zhi Zheng, Yun-Peng Ba et al.· 0 citations
This work proposes Hyper-ES, a subspace-based ES framework that avoids the weakness of ES in full-parameter search while exploiting its strength in low-dimensional optimization, and consistently outperforms GRPO-LoRA while requiring 10% fewer space-consuming gradient updates.
Yuntian Gu, Zhi Zheng, Yun-Peng Ba et al.· 0 citations
These findings position ES as a distinct reasoning post-training paradigm rather than a less effective, memory-efficient alternative to GRPO, and study how hyperparameter design affects the effectiveness of ES, demonstrating that ES requires a smaller population size in a larger LLM.
Yunpeng Ba, Zhi Zheng, Yue Xie et al.· 0 citations
Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to fine-tune larger LLMs...
Zhi Zheng, Rong-Sheng Chen, Yunpeng Ba et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.