Poster: Task-Aware Dynamic Visual Token Pruning for Efficient Vision-Language-Action Inference
Vision-language-action (VLA) models map language instructions and multi-view observations to robot actions, but dense visual-token processing imposes substantial inference overhead. This paper presents TADP, a training-free dynamic visualtoken pruning framework for efficient VLA inference. TADP estimates task-condition...