This paper proposes Hide-and-Seek, a framework that formulates VLA failure detection as a coarsely supervised learning problem that achieves state-of-the-art multi-task failure detection performance with a practical accuracy--timeliness trade-off under conformal prediction, and generalizes well to both seen and unseen...
S. Park, Wendi Li, Changdae Oh et al.· arXiv.org· 8 citations
Humanoid robots are expected to perform diverse human-level tasks in daily environments, many of which require precise regulation of interaction forces. While recent vision-language-action (VLA) models have shown promise for semantic planning and visuomotor control, existing humanoid systems primarily represent actions...
Fu-Kang Liu, Yipu Chen, Jaehwi Jang et al.· 0 citations
MineAmongUs is introduced, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action, and ARIA is proposed, a configurable VLM-agent harness that exposes five cognitive-component ablation axes and opens a new path for embodied VLM-agent alignment research.
Jaewoo Ahn, Junseo Kim, Hyunseo Kim et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.