Skip to content

Author

Haoyang Huang

We have 13 of 45 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and...

Shenghe Zheng, Wen-Bo Li, Ji-Yao Zhang et al. · 0 citations
Preprint Aug 2026

EchoWM: Open and Enterable Omnimodal World Models

We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes,...

Song-Chun Zhang, Yao-Wei Li, Junhao Zhuang et al. · 5 citations · ⚡1

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

Across six real-world scenarios, human raters prefer JoyAI-VL-Interaction over the in-app video-call assistants of Doubao and Gemini by a wide margin, and the first open, vision-driven interaction model released together with its training recipe, data, and complete deployable system.

Ding-Yu Yao, Jun Zhou, Chenxu Yang et al. · 10 citations
Preprint Sep 2026

Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation

Action-conditioned video models require large-scale visual data paired with control signals that are temporally aligned with the resulting scene transitions. Such supervision is difficult to obtain from ordinary real-world video because the actions that caused each visual change are typically unknown. We present a larg...

Hao-Yu Wang, Song-Chun Zhang, Hao-Ran Li et al. · 0 citations
Preprint Sep 2026

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat''when reasoning about the physical world. We argue that this spatial bottleneck stems from a profound dimensional mismatch: while VLMs are trained to interpret 2D projections, true spatial reaso...

Yi-Jun Yang, Shenghe Zheng, Wenbo Li et al. · 0 citations
Preprint Aug 2026

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

This work introduces ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video, and adopts an asynchronous Slow-Fast dual-system architecture to make generative WAMs practical for real-time control.

Xiong-Hao Wu, Yi-Jun Yang, Shi-Long Zhou et al. · 2 citations
Preprint Aug 2026

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

The method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation, and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift.

Yi-Cheng Xiao, Wenxun Dai, Xinran Qin et al. · 5 citations
Jul 2026

Self Gradient Forcing: Native Long Video Extrapolation

Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory.

Junhao Zhuang, Shiyi Zhang, Yuxuan Bian et al. · 5 citations
Jul 2026

Perceptual Flow Matching for Few-Step Generative Modeling

Perceptual Flow Matching supervises flow matching in a perceptual feature space using pretrained perceptual models, which substantially improves the few-step generation capability of flow-matching models, reducing the number of sampling steps from 35-50 to 4-8 while preserving generation quality.

Chuyang Zhao, Yifei Song, Hongfa Wang et al. · 2 citations
Preprint Aug 2026

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds.

Nan Duan, Hao-Yang Huang, Wei-Yang Jin et al. · 2 citations · ⚡1
Preprint Jul 2026

Geo3R: Mitigating Spatial Reasoning Hallucination in Multimodal Large Language Models

Geo3R is proposed, a training-free, plug-and-play framework that incorporates geometric evidence and structured 3D reasoning to mitigate spatial reasoning hallucination and substantially reduces spatial reasoning hallucination across diverse MLLMs without additional training, outperforming existing models and methods.

Mingyu Wang, Weilin Jin, Wenbo Li et al. · 0 citations
Jul 2026

HalluScope: Fine-grained Hallucination Diagnosis for Multimodal Large Language Models

Fine-grained hallucination diagnosis for MLLMs is proposed, a new unified task that jointly performs hallucination detection, classification, and interpretable explanation generation and feedback experiments show that the fine-grained diagnostic explanations produced by the model effectively guide target models to corr...

Weilin Jin, Mingyu Wang, Wenbo Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.