ICI-VLA: In-Context Imitation with Spatiotemporally Aligned Demonstrations for Vision-Language-Action Models
Results indicate that a fixed VLA policy can benefit from conditioning on spatiotemporally aligned demonstrations at test time, and present ICI-VLA, a training and retrieval framework that equips a text-action VLM with few-shot test-time adaptation through in-context demonstrations.