Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

HumanCLAW: Can Vision-Language Models Act Through a Body?

This work introduces HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution in a vision-language model (VLM), and builds HumanCLAW-Bench, a database of 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes.

Siyao Li, Jiawei Gu, Shuai Liu et al. · 1 citation
Preprint Jul 2026

Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention

This work proposes Ms.Forcing, an efficient streaming video generation paradigm that adapts spatial granularity to each state's noise level and introduces Homogeneous-Noise-Level DMD, which assembles each fake video from clean predictions sharing the same source noise level, thereby reducing the mismatch between DMD training sequences and inference-time rollouts.

Zekun Li, Xiaoyan Cong, Hongyu Li et al. · 0 citations