Skip to content
Open access

HDA-MoE: Hybrid Parallelism and Dynamic, Adaptive Scheduling for Mixture-of-Experts with 3D Near-Memory Processing

Sep 2026 · IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems · 0 citations · 49 references
Computer Science

TL;DR

HDA-MoE is presented, a framework that optimizes MoE execution on NMP architectures through hybrid parallel deployment and runtime scheduling and integrates an offline hybrid parallel mapping algorithm with an online dynamic and adaptive scheduling mechanism to reduce communication overhead while improving computation utilization.

Abstract

Mixture-of-Experts (MoE) architectures have become a key technique for scaling Large Language Models (LLMs), enabling high model capacity with reduced computational cost. However, this efficiency comes at the expense of increased memory capacity and bandwidth demands. Recent 3D Near-Memory Processing (NMP) architectures, which vertically integrate memory and compute through hybrid bonding, provide high internal bandwidth and energy efficiency, making them attractive for accelerating MoE inference. Nevertheless, the distributed memory and compute organization of NMP systems introduces new challenges for mapping MoE workloads. Existing parallelization strategies, such as Tensor Parallelism (TP) and Expert Parallelism (EP), suffer from either high communication costs or unbalanced computation utilization, leading to inferior efficiency. In addition, the dynamic routing behavior of MoE models further complicates efficient deployment. To address these challenges, we present HDA-MoE, a framework that optimizes MoE execution on NMP architectures through hybrid parallel deployment and runtime scheduling. HDA-MoE integrates an offline hybrid parallel mapping algorithm with an online dynamic and adaptive scheduling mechanism to reduce communication overhead while improving computation utilization. Experimental results show that HDA-MoE achieves a speedup of 1.1x--3.4x over TP, 1.1x--1.5x over EP, 1.1x--3.7x over the Hybrid TP-EP compute-balanced baseline, and 1.1x--1.3x over HD-MoE. Source code is available at https://github.com/PKU-SEC-Lab/HDA-MoE-TCAD26.

Read PDF

Similar papers

Book Open access Sep 2026

OmniPipe: Efficient, Flexible and Scalable Pipeline Parallelism for Large Model Training

OmniPipe is proposed, a flexible bidirectional multi-pipeline parallelism scheme for unified dense and MoE LLM training that minimizes the pipeline bubble ratio while effectively overlapping EP communication with computation, enabled by the flexible and scalable parallelism scheme of bidirectional pipelines.

Jun Li, Zhi Ma, Shi-Gang Li · 0 citations
#small language model Book Open access Aug 2026

Balancing and Beyond: Communication-Centric Optimizations in Expert Parallelism

EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap.

Jia-Min Cao, Qingxu Li, Yaozhong Liu et al. · 0 citations
Preprint Sep 2026

Weave: Fine-Grained Dynamic SM Scheduling in an MoE Megakernel for Compute-Communication Overlap

Weave is presented, to the authors' knowledge the first MoE overlap system that performs fine-grained dynamic SM scheduling - deciding per layer and per GPU by routing results at runtime, and achieves a 2.89x geometric-mean MoE-layer speedup and a 1.33x geometric-mean end-to-end speedup over five state-of-the-art basel...

Ziyu Huang, Yangjie Zhou, Chen-Hao Zhu et al. · 0 citations
Preprint Aug 2026

NOVA: Technology-Architecture Co-Design of Near-Memory Processing for Attention-SSM-MoE Hybrid LLM Inference

NoVA is proposed, a technology-architecture co-designed NMP system that overcomes both walls and achieves on average 4.5x higher throughput, 69.8% lower end-to-end latency, and 5x better energy efficiency over a GPU baseline, with only 3.9% area overhead and no loss in memory capacity.

In Jung, Jaeha Min, Joo-Young Kim · 0 citations
#machine learning Preprint Sep 2026

Mira: Memory-Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging

Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-constrained, single-GPU systems. However, this benefit is difficult to realize because expert parameters dominate memory, and token-level routing is dynamic, unpredictab...

Sanjali Yadav, Bahar Asgari · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.