Preprint
Jul 2026
HumanCLAW: Can Vision-Language Models Act Through a Body?
This work introduces HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution in a vision-language model (VLM), and builds HumanCLAW-Bench, a database of 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes.
Siyao Li, Jiawei Gu, Shuai Liu et al.
· 1 citation