Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models
Direction-Scale Decomposition (DSD), an action representation that decomposes translation and rotation increments into direction and scale components before tokenization, is introduced, suggesting its potential to mitigate performance degradation when training on large and diverse dataset mixtures.