Skip to content

Predict Before You Deploy: Offline Prediction of Quantization-Induced Task Degradation for World Action Models

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

PreDE (Predict Before You Deploy), a policy-calibrated framework for predicting quantization-induced task degradation from offline action deviations, supports policy-specific behavioral calibration for quantization configuration selection while identifying candidates that require closed-loop evaluation.

Abstract

World action models (WAMs) rely on video-generation backbones, requiring substantial memory and compute for deployment. Post-training quantization reduces memory and can accelerate inference, but bit width, grouping, and quantizer choice define a large configuration space. Identifying configurations that preserve task performance through exhaustive closed-loop evaluation is costly. We propose PreDE (Predict Before You Deploy), a policy-calibrated framework for predicting quantization-induced task degradation from offline action deviations. Using closed-loop outcomes from a small development set, PreDE calibrates two thresholds and accepts, rejects, or defers new configurations using a fixed observation log. Under a within-setting label-ordering hypothesis, the rule issues decisions where all thresholds consistent with the development labels agree. Across five WAMs and four benchmark settings, quantization produces configuration-dependent task losses that cannot be explained by bit width alone or a shared deviation threshold. Across 28 held-out configurations from two policies, PreDE issued 21 decisions before observing closed-loop outcomes (75% coverage), all matching the observed acceptable or degraded labels. Deferred candidates included both acceptable outcomes and a 33-percentage-point loss. In 450 Franka Research 3 trials across two independently fine-tuned policies, all configurations assigned to high-deviation groups before testing showed significant degradation, while low-deviation comparisons showed no statistically significant degradation. On the real robot, W4A4 achieved a 1.37x action-query speedup and approximately 44% lower peak memory. These results support policy-specific behavioral calibration for quantization configuration selection while identifying candidates that require closed-loop evaluation. The code is available at https://github.com/jiuyixu25/PreDE.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models

OnPTQ is proposed, an on-policy framework that calibrates on trajectories visited by the current quantized policy, and derives a Decision--Consequence bound linking behavioral deviation to current policy discrepancy and action-conditioned future-value span.

Wen-Xiao Fan, Jing-Ling Fu, Li-Chen Ma et al. · 0 citations
Preprint Sep 2026

What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling

World action models (WAMs) predict the future alongside actions during training. Due to the heavy computation cost of video denoising, whether the future must still be generated during inference is disputed: Explicit WAMs denoise it into clean frames along with every action chunk, whereas Latent WAMs discard it entirel...

Ren-Ping Zhou, Zan-Lin Ni, Zi-Hao Fan et al. · 0 citations
Preprint Sep 2026

SteerQuant: Steering Quantization Error with Action-Guided Scaling in World-Action Models

World-action models (WAMs) jointly generate future world states and actions through iterative denoising, using shared weights to process heterogeneous semantic streams of video, proprioceptive, and action tokens. Quantization reduces inference cost, but comparable numerical errors in different streams can have markedly...

Yun-Han Wang, Hao-Dong Wang, Zhi-Ming Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning

Quantization-aware distillation (QAD) restores much of the short-form question-answering performance lost to sub-3-bit quantization, yet leaves mathematical and code reasoning substantially impaired. Long generations often degenerate into repetitive loops, exhausting the decoding budget without completing a solution. W...

Yuan-Teng Chen, Zhi-Lei Liu, Peisong Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Scaling Video Generation for Reasoning: At What Cost?

Symbolic state supervision raises the 20M model's frame accuracy from 31.1% to 67.3% at the same training-data budget, suggesting that learning representations of state changes can complement scaling.

Wei-Hang Guo, Xiao-Yu Wu, Yi-Fei Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models

Fine-tuning a Vision-Language-Action (VLA) model for a new deployment environment is expensive, yet most methods apply uniform-capacity adapters to every network region as if every region requires equal adjustment. This paper tests that assumption on five architecturally diverse VLAs (OpenVLA-OFT, $\pi_0$, SmolVLA, DTP...

Shahram Najam Syed, Arthur Jakobsson, Prayuj Sachdev et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.