Sep 2026· Proceedings of the International Conference on Parallel Processing· 0 citations· 29 references
TL;DR
CARE-MoE is proposed, an efficient MoE LLM inference framework comprising two core components that balances expert placement by jointly modeling co-activation correlation and hot–cold drift, preventing overload from correlated experts and enabling low-cost adaptive rebalancing.
Abstract
Mixture-of-Experts (MoE) has become a mainstream architecture for Large Language Models (LLMs) due to its sparse activation mechanism. While distributed inference is promising for deploying LLMs on resource-constrained edge devices, it still faces two critical challenges for MoEs. First, the gating network’s strong bias toward a small subset of hot experts causes severe cross-device load imbalance. Cloud-side methods, such as expert replication, incur prohibitive memory overhead, and global load balancing introduces excessive communication latency in edge networks. Existing edge methods overlook expert co-activation correlations and cannot adapt to hot-cold expert drift induced by shifting user interactions. Second, the All-to-All communication overhead of expert parallelism can dominate the inference latency under bandwidth-constrained edge environments. Existing token dropping or fusion strategies reduce communication at the cost of distorted token feature propagation and degraded long-context accuracy. To address these challenges, we propose CARE-MoE, an efficient MoE LLM inference framework comprising two core components. The Correlation-Aware Expert Placer (CAEP) balances expert placement by jointly modeling co-activation correlation and hot–cold drift, preventing overload from correlated experts and enabling low-cost adaptive rebalancing. The Communication-Aware Expert Router (CAER) exploits semantic equivalence in MoEs to redirect token-expert assignments toward low-latency devices, thereby achieving significant communication reduction with negligible accuracy loss. Experiments across diverse edge environments and MoE models demonstrate that CARE-MoE achieves a 2.1 × to 5.2 × speedup over state-of-the-art baselines.
Sparse expert activation reduces MoE models'computation, yet expert weights can exceed limited device memory. Offloading makes inference feasible on a compact AI appliance but exposes host-to-device transfers to the inference path. We present MoE-CORE, a system that coordinates expert offloading and residency for memor...
Ke Yang, Yong-Ji Gao, Xu-Shi Li et al.· 0 citations
Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory and bandwidth constraints. While speculative decoding and Mixture-of-Experts (MoE) have been proposed to improve inference efficiency, naively combining them often incurs excessive verification overhead and poor exp...
Motivated by the observation that routing patterns are task-specific, MigMoE is proposed, a task-aware expert migration framework that dynamically adjusts expert placement for multi-task expert parallel MoE inference to balance loads across GPUs.
Xu Han, Zi-Nuo Cai, Zhuo-Long Jiang et al.· Proceedings of the Internati...· 0 citations
Mixture-of-Experts (MoE) large language models improve inference efficiency through sparse expert activation, but deployment on resource-constrained devices remains challenging due to the large expert parameter footprint. Expert offloading mitigates this issue by loading experts on demand, yet its effectiveness critica...
Yao Mu, Fa-Hao Chen, Wen-Bin Zhu et al.· Proceedings of the Thirty-Fi...· 1 citation
EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap.
Jia-Min Cao, Qingxu Li, Yaozhong Liu et al.· Conference on Applications,...· 0 citations
A ReRAM near-memory architecture that keeps expert weights resident behind high-bandwidth local reads and recovers occupancy with bounded core-local multicast pooling, coactivation-aware placement, and load-aware fetch, and sizes each communication level from induced demand is presented.
Kun-Ming Shao, Ming Zeng, Xin Yuan et al.· 0 citations
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.