Skip to content

Author

Heng-Shuang Zhao

We have 4 of 45 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Sep 2026

Enhancing Local Cognition of CLIP for Training-Free Open Vocabulary Semantic Segmentation.

CLIP, as a vision-language model, has significantly advanced Open-Vocabulary Semantic Segmentation (OVSS) with its zero-shot generalization. Despite its success, its application to OVSS is limited due to its initial image-level alignment training, which affects its performance in tasks requiring detailed local context....

Tong Shao, Zhuo-Tao Tian, Yun-Yang Mo et al. · 0 citations
Preprint Aug 2026

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically eval...

Kai Ding, Xi Chen, Minghong Cai et al. · 1 citation
Preprint Aug 2026

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-d...

Senqiao Yang, Chengyao Wang, Yuxin Chen et al. · 2 citations
Preprint Aug 2026

OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing

The OmniEdit-Bench provides a comprehensive and reliable testbed for evaluating instruction-based video editing and offers insights into future research directions, including an accuracy-aware penalty mechanism that conditions other scores on accuracy, preventing visually plausible but incorrect edits from receiving in...

Chenxuan Miao, Yutong Feng, Yi Lu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.