Skip to content

VLAQuantBench: Closed-Loop Evaluation of Post-Training Quantization for Vision-Language-Action Models

Sep 2026 · 0 citations · 36 references
Computer Science

TL;DR

These results establish recipe-dependent interactions and identify concrete precision assignments, rather than universal layer-sensitivity rules, rather than universal layer-sensitivity rules for vision-language-action models.

Abstract

Post-training quantization reduces the memory requirements of vision-language-action (VLA) models, but precision selection must account for the interaction between layer scope, numerical format, and calibration. We introduce \textbf{VLAQuantBench}, a controlled evaluation with 409 runs and 94,574 simulation episodes: four models on LIBERO, with X-VLA additionally evaluated on three simulation benchmark families. Under uncalibrated W4A4 round-to-nearest quantization, expanding a $\pi_{0.5}$ action-head subset from 126 to 167 layers raises success from 7.0\% to 70.5\%. Fixed-observation replay confirms a corresponding numerical recovery. Two-episode calibration removes the severe joint failures in the tested subsets, whereas the same smoothing-and-clipping recipe lowers $\pi_0$ success and does not recover OpenVLA-OFT end-to-end. For OpenVLA-OFT, protecting one 28,672-parameter output projection instead restores near-baseline success: the remaining 441 eligible linear layers retain W3 on LIBERO-Long or eight-bit activations across all four suites. Task-clustered intervals support the large failure and recovery contrasts. These results establish recipe-dependent interactions and identify concrete precision assignments, rather than universal layer-sensitivity rules. Real-kernel and physical-robot measurements complement the accuracy analysis. Code, configurations, and episode records are publicly available at https://github.com/jiuyixu25/VLAQuantBench.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Validating, Not Sampling: Region-Level Robustness of Vision-Language and Vision-Language-Action Models

Vision-language models (VLMs) and vision-language-action models (VLAs) are increasingly deployed in real-world applications. There, a small perturbation to the recorded camera image may change a decision significantly. However, existing benchmarks for these models only sample perturbations, which does not guarantee the...

Bogdan Aron, Christopher Brix, Benedikt Brückner et al. · 0 citations
Preprint Sep 2026

IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

Vision-language-action (VLA) policies leverage pretrained vision-language backbones to achieve strong cross-task generalization. A leading design couples this backbone with a dedicated continuous action head trained via diffusion or flow matching. However, such heads rely on iterative multi-step sampling, for example 1...

Kian Hosseinkhani, Qin-He Peng, George Shramko et al. · 1 citation
Preprint Sep 2026

FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

Recurrent Action Memory (RAM) is proposed, a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking.

Kemal Oksuz, Alexandru Buburuzan, Yu-Han Yao et al. · 0 citations
Preprint Sep 2026

Recovering Aggressively Pruned Vision-Language-Action Models with Offline Hidden-State Distillation

Vision-language-action (VLA) models let robots follow language instructions, but their language backbones of several billion parameters are the main obstacle to running them on robot hardware. Structured pruning reduces that backbone, and removing 63% of it from OpenVLA-OFT drops LIBERO-Long success from 93.2% to 0.8%....

Chiyoung Kim, S. Choi, Minhyeok Lee · 0 citations
#artificial intelligence Preprint Oct 2026

VLA-ZO: Fast Zeroth-Order Adaptation for Vision-Language-Action Models

Adapting vision-language-action (VLA) models to deployment-time distribution shifts is important for reliable robotic operation, but conventional first-order adaptation can exceed the memory budget of inference-oriented deployment platforms. Zeroth-order (ZO) optimization offers a forward-only alternative with inferenc...

Jaemin Kim, Jiahn Kim, Taesik Gong · 0 citations
#machine learning Preprint Sep 2026

Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models

Fine-tuning a Vision-Language-Action (VLA) model for a new deployment environment is expensive, yet most methods apply uniform-capacity adapters to every network region as if every region requires equal adjustment. This paper tests that assumption on five architecturally diverse VLAs (OpenVLA-OFT, $\pi_0$, SmolVLA, DTP...

Shahram Najam Syed, Arthur Jakobsson, Prayuj Sachdev et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.