HumanCLAW: Can Vision-Language Models Act Through a Body?
This work introduces HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution in a vision-language model (VLM), and builds HumanCLAW-Bench, a database of 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes.