Preprint
Sep 2026
ICI-VLA: In-Context Imitation with Spatiotemporally Aligned Demonstrations for Vision-Language-Action Models
Results indicate that a fixed VLA policy can benefit from conditioning on spatiotemporally aligned demonstrations at test time, and present ICI-VLA, a training and retrieval framework that equips a text-action VLM with few-shot test-time adaptation through in-context demonstrations.
Song-Hua Yang, Zi-Yu Liu, Xue-Tao Li et al.
· 0 citations