Vision-language-action (VLA) models have shown strong potential as generalist robot policies, but adapting them to unseen tasks often requires costly parameter updates. Recent work such as RICL introduces in-context adaptability by retrieving expert demonstrations based on the current VLA observation and providing them...
Zi-Xuan Liu, J. Koster, Zizhan Zheng et al.· 0 citations
A teacher-student knowledge distillation framework in which large VLMs fine-tuned for fire understanding can be distilled into lightweight students is developed, which provides broader guidance for deploying domain-specialized VLMs in resource-constrained, safety-critical settings.
Mohammad Kazzazi, Zi-Xuan Liu, S. Khajavi· 0 citations
The results suggest that the accuracy--cost trade-off of long-term agent memory depends not only on what information is retained, but also on the granularity at which it is consolidated, and that the accuracy--cost trade-off of long-term agent memory depends not only on what information is retained, but also on the gra...
Dongfang Li, Zi-Xuan Liu, Junmai Wang et al.· 3 citations
CAER introduces a span-grounded evidence router that transforms claim representations into soft textual queries and retrieves corresponding evidence from frozen visual tokens, enabling fine-grained conflict estimation and design a dual-prefix expert routing mechanism that learns separate experts for visually supported...
Zi-Xuan Liu, Juntao Cai, Xiaoxu Cai et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.