Mixture-of-Experts (MoE) has increasingly become a mainstream approach for scaling large language models, as it expands model capacity while keeping computation cost nearly constant. Training large-scale MoE models relies on Expert Parallelism (EP), which distributes expert replicas across GPUs and exchanges tokens thr...
Production LLM serving multiplexes hundreds of heterogeneous models on shared clusters, exposing three challenges that existing systems fail to address simultaneously: unpredictable bursts, power-law application popularity, and heterogeneous yet complementary resource demands. We present Janus, a Service-Engine co-desi...
Tian-Bao Zhou, Yi Wang, Yu Zhou et al.· Proceedings of the ACM SIGOP...· 0 citations
It is shown that pipeline parallelism (PP), long overlooked because it offers little decode-latency advantage, can reduce JCT by providing a more favorable balance between prefill and decode efficiency, and PipeSwift is built, an optimized open-source pipeline-parallel runtime integrated with a tailored micro-batch par...
Shiju Wang, Fei Ren, Fang-Cheng Fu et al.· 0 citations
This paper proposes DPP, which transforms PP granularity from a static design choice into a workload-adaptive optimization space over packed, split, and hybrid chunks and further introduces a new coupling between heterogeneous pipeline scheduling and gradient checkpointing.
Shiju Wang, Yujie Wang, Fang-Cheng Fu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.