Skip to content

Author

Pei-Yu Zang

We have 5 of 13 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

SubZero+: Memory-Efficient Adaptive Zeroth-Order LLM Fine-Tuning in Random Subspaces

Zeroth-order (ZO) optimization with SGD in random subspaces enables memory-efficient fine-tuning of large language models without backpropagation. However, high gradient estimation noise fundamentally undermines adaptive optimizers like Adam. We propose SubZero+, which achieves practical adaptive ZO optimization throug...

Zi-Ming Yu, Shu-Yao Xiao, Xingyu Zhao et al. · 0 citations
Jul 2026

FlashPDE: A Drop-In Fused Triton Operator Library for Neural PDE Solvers

FlashPDE provides a hardware-efficient execution layer that bridges differentiable PDE solvers and GPU-optimized numerical computation within the PyTorch ecosystem, while maintaining numerical agreement with PyTorch finite-difference references.

Pei-Yu Zang, Bosen Xie, Ruoxi Xu et al. · 1 citation
Preprint Jul 2026

KernelGenBench: Can LLMs and Agents Write Efficient Kernels Across Operator Sources and Hardware Platforms?

KernelGenBench is presented, the first unified multi-source and multi-chip infrastructure for evaluating LLM- and agent-generated Triton kernels and establishes operator source, hardware platform, and agentic scaffold as distinct dimensions of kernel-generation capability, and shows that success in a familiar source-ha...

Pei-Yu Zang, Jian-Hang Tao, Jia-Ling Zhang et al. · 0 citations
Jul 2026

KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation

Large language models (LLMs) have significantly increased the demand for efficient accelerator kernels, but kernel development remains a highly specialized and labor-intensive task. The recent rise of LLMs and agentic frameworks offers a promising pathway toward automatic kernel generation. However, despite rapid progr...

Pei-Yu Zang, Jian-Hang Tao, Jia-Ling Zhang et al. · 1 citation
#machine learning Preprint Aug 2026

SubZero+: Efficient Zeroth-Order LLM Fine-Tuning via Large Learning Rates

Experiments show that SubZero+, an improved SubZero framework that improves stability in three complementary ways, consistently outperforms prior ZO baselines, enlarges the stable learning-rate range, and narrows the gap to first-order methods with minimal extra memory overhead.

Ziming Yu, Shu-Yao Xiao, Xingyu Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.