Skip to content

Latency-Tolerant Cloud-Edge Collaborative Vision-Language-Action Models via Emergent Representational Specialization

Aug 2026 · 1 citation · 27 references
Computer Science Engineering

TL;DR

CloudEdgeVLA is introduced, a cloud-edge policy that treats temporal misalignment as a representation-learning problem and offers a practical route to scalable VLA deployment in which cloud models can grow while edge computation remains lightweight and responsive.

Abstract

Deploying billion-parameter Vision-Language-Action (VLA) policies on mobile robots creates a systems conflict: semantic reasoning benefits from cloud GPUs, whereas closed-loop control must respond locally despite network delay and jitter. Existing hierarchical and asynchronous policies improve throughput, but their slow-path representations can still arrive stale or require explicit scheduling and delay cues. We introduce CloudEdgeVLA, a cloud-edge policy that treats temporal misalignment as a representation-learning problem. A cloud VLA encodes delayed observations into slowly varying task features, while a lightweight edge head combines the latest available cloud feature with current local vision. During training, current and randomly delayed frames are paired with the same current action target in fresh and stale paths. This objective encourages the cloud representation to preserve task-level information while the edge path supplies state-sensitive corrections, driving emergent specialization. Across four LIBERO suites, CloudEdgeVLA retains 63.8-78.0% success with a 40-step uniform-delay window, whereas VLASH reaches at most 6.4% and the evaluated single-path baselines at most 3.0%. By removing blocking synchronization from the control loop, the design offers a practical route to scalable VLA deployment in which cloud models can grow while edge computation remains lightweight and responsive.

View source

Similar papers

Preprint Aug 2026

Risk-Adaptive Edge--Cloud Visual Reasoning for Communication-Efficient Autonomous Driving

A risk-adaptive edge-cloud architecture in which onboard traffic assessment determines when cloud reasoning is requested is presented, in which onboard traffic assessment served as a practical trigger for selective VLM inference in these experiments.

Meng Ma, Shu-Yang Li, Nai-Gang Wang et al. · 0 citations
Preprint Sep 2026

EdgeVLN: Runtime-Aware Deployment Ready Quantized Vision Language Navigation Model

Vision-language navigation (VLN) models perform well but target compute-rich platforms, limiting deployment on memory- and power-constrained robotic edge devices. Compression alone does not establish whether a VLN model fits the memory, latency, and energy budgets of an edge platform while preserving navigation behavio...

Rithvik Jonna, Man Namgung, Aakash Gurram et al. · 0 citations
Conference Aug 2026

VTP: A Task-Aware Transmission Protocol for Edge Vision-Language-Action Robotic Systems

Vision-Language-Action (VLA) models are increasingly offloaded to edge servers, making visual transmission critical for continuous robotic execution. Under packet-loss conditions, delayed or incomplete visual delivery may postpone the generation of subsequent action chunks. However, visual transmission in edge-assisted...

Richeng Huo, Yan-Sen Wang, Ye Wu et al. · 0 citations
#machine learning Preprint Sep 2026

Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs

This work proposes a framework that exposes backbone depth V, action expert depth A, and denoising steps $D$ as three jointly configurable compute axes in a VLA, and introduces a KV Cache synthesis mechanism that manages the missing keys and values of the skipped backbone layers, allowing the action expert to exit deep...

Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro et al. · 0 citations
Preprint Aug 2026

SparkVLA: Stop-Aware Hierarchical VLA with Adaptive Action Chunking for Long-Horizon Manipulation

SparkVLA is presented, a stop-aware hierarchical VLA that resolves this mutual dependency by formulating both decisions as a single ranking: Stop competes against every action-prefix length in a unified candidate set, and the system selects the highest-scoring option, eliminating threshold tuning and requiring only off...

Xun-Yao Lei, Ren-Jun Wu, Tianlin Huo et al. · 1 citation
Preprint Sep 2026

TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation

Reactive vision-language-action (VLA) policies suffer from task-state aliasing in long-horizon manipulation, where identical multimodal inputs call for distinct, context-dependent actions. Given that pretrained VLAs already possess rich control primitives to express diverse behaviors, we hypothesize that the execution...

Heng-Yan Liu, Wen-Lve Zhou, Bo Yue et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.