Sparse expert activation reduces MoE models'computation, yet expert weights can exceed limited device memory. Offloading makes inference feasible on a compact AI appliance but exposes host-to-device transfers to the inference path. We present MoE-CORE, a system that coordinates expert offloading and residency for memor...
Ke Yang, Yong-Ji Gao, Xu-Shi Li et al.· 0 citations
Chunked prefill pipeline parallelism (CPP) is a key technique for LLM inference. However, equal-size chunks exhibit imbalanced latency, as later chunks attend longer prefix KV caches and incur higher attention costs, leading to pipeline bubbles. Existing approaches mitigate this imbalance through dynamic chunk resizing...
Yan Shi, Xiao-Chao Wang, Jing-Chun Gao et al.· 0 citations
This work forms kernel optimization as a progressive cross-layer diagnosis problem that links runtime symptoms to IR structure and compiler behavior before rewriting source, and presents a compiler-grounded and hierarchical optimization framework for Triton kernels.