Skip to content

Author

Jian-Hua Han

We have 4 of 87 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these metrics evaluate perceptual dimensions sep...

Yin-Ming Huang, Shu-Yuan Tu, Xi Yan et al. · 4 citations
Preprint Aug 2026

4D-WAM: 4D Consistent World Modeling for Autonomous Driving

Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and trajectory planning. However, existing WAMs are typically trained with video data, which is only 2D projections of the underlying 4D driving scene. Consequently, WAMs...

Jiacheng Fu, Yibo Yuan, Meng Tian et al. · 0 citations
Preprint Aug 2026

SUV: Future Scene Understanding as Video Generation for End-to-End Driving

End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalability. Can video generation instead provide a shared predictor? We introduce SUV, a unified end-to-end driving framework that casts future Scene...

Yibo Yuan, Jiacheng Fu, Jiangtong Zhu et al. · 0 citations
Preprint Jul 2026

Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools

Recent multimodal large language models (MLLMs) have made remarkable progress on fine-grained perception tasks under the"Thinking with Images"(TwI) paradigm by iteratively performing various visual tool operations. However, this paradigm relies heavily on frequent external tool calls and repeated image re-encoding, whi...

Xiuwei Chen, Quanlin Chen, Wentao Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.