Skip to content

Similar papers

Preprint Sep 2026

A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies

Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal inputs, including semantic instructions, visual observations of the scene, and proprioceptive observations, to robot actions. Most state-of-the-art models predict actions in the end-effector pose space as sequences of action chu...

Mathilde Kappel, Clémence Grislain, Mohamed Chetouani et al. · 0 citations
Preprint Sep 2026

Recovering Aggressively Pruned Vision-Language-Action Models with Offline Hidden-State Distillation

Vision-language-action (VLA) models let robots follow language instructions, but their language backbones of several billion parameters are the main obstacle to running them on robot hardware. Structured pruning reduces that backbone, and removing 63% of it from OpenVLA-OFT drops LIBERO-Long success from 93.2% to 0.8%....

Chiyoung Kim, S. Choi, Minhyeok Lee · 0 citations
Preprint Aug 2026

TemporalFlow-VLA: Learning Physically Grounded Execution History for Long-Horizon Robot Manipulation

TemporalFlow-VLA provides a compact, physically grounded interface for exploiting ordered execution history without explicit motion estimation or geometric processing at deployment, and shows its clearest advantage over prior methods on longer-horizon, multi-stage manipulation.

Jia-Rui Yang, Ye-Hao Lu, Yu-Ning Su et al. · 1 citation
#machine learning Preprint Sep 2026

Reinforcement Learning for Real-Time Vision-Language-Action Policies

Reinforcement learning fine-tuning on top of large, pretrained Vision-Language-Action (VLA) models offers promise for highly reliable robot deployment. However, because of their scale, modern VLA models suffer from high inference latency, so the observation used to select an action is often stale by execution time, cre...

Perry Dong, Kuo-Han Hung, D. Sadigh et al. · 0 citations
#machine learning Preprint Sep 2026

Efficient Vision-Language-Action Management and Serving for Robot Factories

Vision-Language-Action (VLA) models show high robotic manipulation capabilities via a two-stage design: a Vision-Language Model (VLM) stage followed by an Action Diffusion Transformer (ADiT) stage. Since robots must meet strict Service-Level Objectives (SLOs) for safety, VLA inference is inherently latency-critical. Me...

Dionysios Adamopoulos, Nattapol Chanpaisit, Basel Fakhri et al. · 0 citations
#small language model Preprint Sep 2026

WorldContact: A Contact-Centric World Model for Scalable Robot Learning

Adapting robots to new objects and tasks requires interaction experience that can be costly to obtain. We present WorldContact, a contact-centric world model for deformable-object manipulation, constructed from a limited set of high-quality trajectories to generate additional training data efficiently. It predicts obje...

Caoliwen Wang, Meng-Di Wang, Heng Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.