FedMVLA is proposed, a modality-decoupled FL framework for privacy-preserving embodied intelligence in 6G networks that incorporates three mechanisms: modality-aware federated aggregation (MAFA), modality-aware privacy allocation (MAPA), and modality-aware communication compression (MACO), together with a modality-sliced transport design that routes the precision-critical action stream through a protected ultra-reliable low-latency slice.
Abstract
Sixth-generation (6G) wireless networks are expected to provide a key infrastructure for large-scale embodied intelligence, where heterogeneous robots collaborate through low-latency connectivity, edge intelligence, and distributed sensing. Vision-language-action (VLA) models offer a foundation by integrating visual perception, language understanding, and action generation into a unified closed-loop policy. However, training and adapting VLA models to distributed robotic agents introduce challenges in privacy protection, communication efficiency, and model heterogeneity. Existing federated learning (FL) methods overlook the intrinsic differences among vision, language, and action pathways in parameter scale, privacy exposure, update dynamics, and tolerance to compression or perturbation. To address this issue, this article proposes FedMVLA, a modality-decoupled FL framework for privacy-preserving embodied intelligence in 6G networks. FedMVLA incorporates three mechanisms: modality-aware federated aggregation (MAFA), modality-aware privacy allocation (MAPA), and modality-aware communication compression (MACO), together with a modality-sliced transport design that routes the precision-critical action stream through a protected ultra-reliable low-latency slice. A case study on federated robotic manipulation over the Third Generation Partnership Project (3GPP)-based wireless substrate, covering fading, co-channel interference, and malicious jamming, shows that FedMVLA achieves an 84.8% task success rate, exceeds FedAvg by 22.2 percentage points, sustains a widening margin when scaling to 128 clients across eight cells, and reduces the schedule-averaged per-client uplink model-update payload by 95.6% (approximately 96%), while keeping the 95th percentile (p95) of the round-critical uplink completion time near 1.5s.
FedMARL-LTI is presented, a federated multi-agent reinforcement learning framework whose architecture answers both pressures with a single decision: each organization’s threat intelligence is shared only as a differentially private 768-dimensional semantic embedding, never as raw data.
Federated learning offers a natural way for multiple robots to jointly improve manipulation policies without requiring centralized access to training demonstrations. However, non-IID task and environment distributions can induce representation drift and mutually incompatible robot-policy updates, making naive parameter...
Biprodip Pal, Kaushik Roy, Yan-Ming Zhu et al.· 0 citations
This article presents the first dedicated and comprehensive survey of DRL for Open AI-RAN, and reviews the foundations of model-free, model-based, offline, safe, multi-agent, federated, and transfer learning, and provides an O-RAN-aware framework for formulating RAN control problems through states, observations, action...
Jie Lu, Peihao Yan, Qijun Wang et al.· 0 citations
Sixth generation (6G) wireless networks aim to move beyond connected devices toward a “connected cognition” model in which the network understands intent, reasons about constraints, and executes actions autonomously. This article proposes an Agentic-Native 6G architecture embedding intelligence directly into network fu...
This paper addresses the challenges of dynamic resource allocation and scheduling in 6G in-X subnetworks supporting applications with heterogeneous characteristics by proposing a novel framework that combines Multi-Agent Reinforcement Learning (MARL), Federated Learning (FL), and Explainable AI (XAI).
Experimental results on multiple datasets show that the proposed DP-aided FedSFR outperforms DP-enabled FedAvg in training stability and image reconstruction quality in heterogeneous wireless systems.