Open access
Aug 2026
VLM-guided retrieval augmented generation (RAG) for robot action prediction
A VLM-guided retrieval-augmented generation (RAG) framework that crops task-relevant image regions, retrieves examples from a structured local database, re-ranks candidates, and expands the search when needed enables accurate, data-efficient adaptation to new device types without fine-tuning, at the cost of increased and variable latency.
Boris Kuster, Nikola Marić, Fatima Aziz et al.
· Frontiers in Neurorobotics · 0 citations