Validating, Not Sampling: Region-Level Robustness of Vision-Language and Vision-Language-Action Models
Vision-language models (VLMs) and vision-language-action models (VLAs) are increasingly deployed in real-world applications. There, a small perturbation to the recorded camera image may change a decision significantly. However, existing benchmarks for these models only sample perturbations, which does not guarantee the...