Skip to content
Conference

VTP: A Task-Aware Transmission Protocol for Edge Vision-Language-Action Robotic Systems

Aug 2026 · 2026 IEEE/CIC International Conference on Communications in China (ICCC) · pp. 617-622 · 0 citations · 19 references

Abstract

Vision-Language-Action (VLA) models are increasingly offloaded to edge servers, making visual transmission critical for continuous robotic execution. Under packet-loss conditions, delayed or incomplete visual delivery may postpone the generation of subsequent action chunks. However, visual transmission in edge-assisted VLA systems is task-dependent: large-motion phases are relatively tolerant of incomplete historical visual information, whereas fine-grained manipulation phases require more complete visual inputs. Meanwhile, different visual frames contribute unequally to VLA inference results. Conventional transport protocols treat packets uniformly and do not consider VLA task requirements. Based on these observations, we propose the Task-Aware VLA Transmission Protocol (VTP), a customized transmission protocol for edgeassisted VLA systems. VTP estimates task phase from inter-frame visual activity and adapts packet-level transmission according to task phase and frame importance. Specifically, VTP prioritizes packets of critical anchor frames while selectively aborting delayed recovery of prior-frame packets. At the edge server, VTP reconstructs unavailable prior frames through interpolation to maintain continuous visual inputs for the VLA model. Extensive experiments show that VTP reduces visual delivery delay and improves deadline-bounded task success, achieving an 85.0% task success rate under a 10% packet loss rate.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.