Skip to content
Preprint

Pruning-Aware Multi-Cluster Co-Inference for Large AI Models in AI-RANs

Aug 2026 · 0 citations · 38 references
Computer Science

TL;DR

A multi-cluster LAIM co-inference framework, where an edge server equipped with multiple graphics processing units (GPUs) coordinates multiple user clusters to execute inference tasks collaboratively, significantly outperforms existing benchmark schemes, achieving superior inference accuracy and resource efficiency in multi-cluster edge intelligence networks.

Abstract

The increasing scale and computational demands of large artificial intelligence models (LAIMs) present significant challenges for efficient inference in resource-constrained distributed environments. In this paper, we propose a multi-cluster LAIM co-inference framework, where an edge server equipped with multiple graphics processing units (GPUs) coordinates multiple user clusters to execute inference tasks collaboratively. Within each cluster, devices capture data from diverse perspectives and employ lightweight on-device LAIMs to extract local features. These features are then transmitted to the edge server, where they are aggregated and fused to generate a more accurate inference outcome. To reveal the fundamental trade-off between model pruning and collaborative inference performance, we develop a theoretical framework that characterizes the impact of pruning ratios and device contributions using rate-distortion theory and partial information decomposition. Based on this analysis, we formulate a joint optimization problem that determines the model pruning ratio, the task scheduling strategy, the bandwidth allocation, and the transmission power, with the goal of minimizing the inference distortion while satisfying the constraints of latency, energy consumption, and server capacity. Extensive simulation results demonstrate that the proposed framework significantly outperforms existing benchmark schemes, achieving superior inference accuracy and resource efficiency in multi-cluster edge intelligence networks.

View source

Similar papers

Preprint Aug 2026

Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts

MOSAIC is developed, which formulates model architecture and systems co-design as an optimization problem that couples a predictive scaling law with a calibrated performance model that estimates Model FLOPs Utilization (MFU), communication cost, memory footprint, and the best parallel layout.

Soumajyoti Sarkar, Yu-Xin Tang, Sheng Zha · 0 citations
#edge computing Preprint Sep 2026

Iapetus: Content-Aware Hierarchical Scheduling for Collaborative ViT Inference in LEO Satellite Networks

sys is presented, a content-aware hierarchical scheduler that screens constellation-wide options to retain a bounded candidate set, then refines each candidate into a complete token compression and layer offloading trajectory using quality prediction and joint planning and balances per-task latency, energy, and quality...

Yan Chen, Yun-Xiang Zhang, Guan-Jun Jiang et al. · 0 citations
Sep 2026

Similarity-Aware ML Model Selection for Network-Device Collaboration

The integration of artificial intelligence and machine learning (AI/ML) into 5G-Advanced and emerging 6G systems introduces significant challenges in efficiently managing model distribution between network infrastructure and user equipment (UE). In particular, conventional per-model transfer and activation can incur su...

Hojin Kim, Osvaldo Gonsa · 0 citations
#graph neural networks Open access Sep 2026

Synergizing large and small models in cloud-edge continuum: a spatiotemporal hypergraph approach for dynamic offloading

This paper proposes a spatiotemporal hypergraph-driven framework integrating high-order topological feature extraction with dynamic resource modeling, and introduces dynamic hypergraph sequences to naturally encompass local conflict domains, mitigating the topological blind spots and “over-smoothing” issues inherent in...

Kun Ding, Xi-Wen Qiu, Nian-Feng Weng et al. · 0 citations
Open access 2026

DABO: Difficulty-Aware Binary Offloading for Collaborative Large-Small Model Inference

DABO is proposed, a calibration-aware binary offloading method for collaborative large–small model inference that maintains competitive end-to-end accuracy while processing an average of 83.72% of requests at the edge.

Chen Zhu, Yi-Ming Su, Chenwenjie Mao et al. · 0 citations
Conference Aug 2026

Computing Assignment and Radio Resource Allocation for Edge-End Collaborative Inference in Mobile Edge Computing

The growing usage of AI and large models has driven increasing demand for edge-end collaborative inference in mobile edge computing. However, inefficient resource utilization still prevents effective support for such services, especially reflected in difficulty of coordinating numerous computation requests with limited...

Hao Luo, Hui Tian, Hui-Qing Ao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.