Skip to content
Preprint

Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models

Sep 2026 · 0 citations · 19 references
Computer Science

TL;DR

Direction-Scale Decomposition (DSD), an action representation that decomposes translation and rotation increments into direction and scale components before tokenization, is introduced, suggesting its potential to mitigate performance degradation when training on large and diverse dataset mixtures.

Abstract

Action representation plays a central role in discrete-token vision-language-action (VLA) learning but remains underexamined. Under conventional pose-increment representations, action tokens are sensitive to execution speed and dataset-specific normalization, potentially obscuring geometric structure shared across demonstrations and datasets. We introduce Direction-Scale Decomposition (DSD), an action representation that decomposes translation and rotation increments into direction and scale components before tokenization. DSD isolates motion direction while retaining magnitudes in separate scale channels. We evaluate DSD with uniform binning (BIN) and BEAST, a B-spline-based tokenizer, in simulation and real-world manipulation under both single-dataset and mixed-dataset training. On LIBERO, DSD improves average success rates with both tokenizers. On SimplerEnv, DSD-BIN outperforms BIN by 10.3 percentage points in overall success rate under mixed-dataset training. Real-robot experiments further show gains both with and without robotics pretraining. These results support DSD as an effective action representation for discrete-token VLA models and suggest its potential to mitigate performance degradation when training on large and diverse dataset mixtures. Our project page with additional resources is available at https://vla-dsd.github.io/

View source

Similar papers

#artificial intelligence Preprint Sep 2026

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet sma...

Shijie Lian, Bin Yu, Zhao-Long Shen et al. · 1 citation
Preprint Aug 2026

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-d...

Sen-Qiao Yang, Cheng-Yao Wang, Yuxin Chen et al. · 4 citations
Preprint Aug 2026

Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models

SALT is introduced, a Semantically ALigned action Tokenizer that augments a VQ-VAE-style tokenizer with an auxiliary objective requiring a frozen vision-language model to recover the episode instruction from quantized action latents to substantially improve language-conditioned control.

Wen-Jie Li, Yash Jangir, Ignacy Stepka et al. · 0 citations
Preprint Oct 2026

UniWAM: Unified World-Action Model

Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semant...

Jia-Yi Chen, Wen-Xuan Song, Jing-Bo Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models

Time-Frequency Geometric Cross-Attention (TFGCA), a drop-in module repairing both blind spots in vision-language-action policies, uses a per-dimension learnable stationary wavelet transform to decompose the action chunk into time-frequency tokens, and each time token retrieves information from them via a cross-attentio...

Sheng-Ye Dong, Hao-Chen Niu, Hao Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.