Open access
Aug 2026
Action- and Language-Conditioned Video Assessment for Embodied Control
ALVA (Action- and Language-Conditioned Video Assessment), a trajectory evaluator that conditions its assessment on visual observations, the executed action sequence, and the natural language instruction, provides more effective feedback than the evaluated static image and embedding-based visual baselines and reduces the performance gap to a ground-truth oracle.
Hwanhee Kim, Jaehyun Jang, Seung-Min Cha et al.
· Italian National Conference... · 0 citations