Skip to content
#small language model Book Open access

CARE-MoE: Correlation-Aware Expert Placement and Semantic Equivalence Routing for MoE LLM Inference on Edge Devices

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 29 references

TL;DR

CARE-MoE is proposed, an efficient MoE LLM inference framework comprising two core components that balances expert placement by jointly modeling co-activation correlation and hot–cold drift, preventing overload from correlated experts and enabling low-cost adaptive rebalancing.

Abstract

Mixture-of-Experts (MoE) has become a mainstream architecture for Large Language Models (LLMs) due to its sparse activation mechanism. While distributed inference is promising for deploying LLMs on resource-constrained edge devices, it still faces two critical challenges for MoEs. First, the gating network’s strong bias toward a small subset of hot experts causes severe cross-device load imbalance. Cloud-side methods, such as expert replication, incur prohibitive memory overhead, and global load balancing introduces excessive communication latency in edge networks. Existing edge methods overlook expert co-activation correlations and cannot adapt to hot-cold expert drift induced by shifting user interactions. Second, the All-to-All communication overhead of expert parallelism can dominate the inference latency under bandwidth-constrained edge environments. Existing token dropping or fusion strategies reduce communication at the cost of distorted token feature propagation and degraded long-context accuracy. To address these challenges, we propose CARE-MoE, an efficient MoE LLM inference framework comprising two core components. The Correlation-Aware Expert Placer (CAEP) balances expert placement by jointly modeling co-activation correlation and hot–cold drift, preventing overload from correlated experts and enabling low-cost adaptive rebalancing. The Communication-Aware Expert Router (CAER) exploits semantic equivalence in MoEs to redirect token-expert assignments toward low-latency devices, thereby achieving significant communication reduction with negligible accuracy loss. Experiments across diverse edge environments and MoE models demonstrate that CARE-MoE achieves a 2.1 × to 5.2 × speedup over state-of-the-art baselines.

Read PDF

Similar papers

Preprint Oct 2026

MoE-CORE: Coordinated Expert Offloading and Residency for Memory-Constrained MoE Inference

Sparse expert activation reduces MoE models'computation, yet expert weights can exceed limited device memory. Offloading makes inference feasible on a compact AI appliance but exposes host-to-device transfers to the inference path. We present MoE-CORE, a system that coordinates expert offloading and residency for memor...

Ke Yang, Yong-Ji Gao, Xu-Shi Li et al. · 0 citations
#artificial intelligence Preprint Aug 2026

S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices

Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory and bandwidth constraints. While speculative decoding and Mixture-of-Experts (MoE) have been proposed to improve inference efficiency, naively combining them often incurs excessive verification overhead and poor exp...

Hao-Chen Huang, Sheng-Xuan Qiu, Meng Li · 2 citations
Book Open access Sep 2026

MigMoE: Task-Aware Expert Migration for Faster and More Balanced Expert-Parallel MoE Inference

Motivated by the observation that routing patterns are task-specific, MigMoE is proposed, a task-aware expert migration framework that dynamically adjusts expert placement for multi-task expert parallel MoE inference to balance loads across GPUs.

Xu Han, Zi-Nuo Cai, Zhuo-Long Jiang et al. · 0 citations
Conference Open access Sep 2026

DoMoE: Domain-Aware Semantic Expert Prediction for Efficient MoE Inference Under Expert Offloading

Mixture-of-Experts (MoE) large language models improve inference efficiency through sparse expert activation, but deployment on resource-constrained devices remains challenging due to the large expert parameter footprint. Expert offloading mitigates this issue by loading experts on demand, yet its effectiveness critica...

Yao Mu, Fa-Hao Chen, Wen-Bin Zhu et al. · 1 citation
#small language model Book Open access Aug 2026

Balancing and Beyond: Communication-Centric Optimizations in Expert Parallelism

EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap.

Jia-Min Cao, Qingxu Li, Yaozhong Liu et al. · 0 citations
Preprint Aug 2026

MoE Expert Execution in Disaggregated LLM Serving with a High-Bandwidth ReRAM Near-Memory Architecture

A ReRAM near-memory architecture that keeps expert weights resident behind high-bandwidth local reads and recovers occupancy with bounded core-local multicast pooling, coactivation-aware placement, and load-aware fetch, and sizes each communication level from induced demand is presented.

Kun-Ming Shao, Ming Zeng, Xin Yuan et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.