Skip to content

Author

Yinqiang Zheng

We have 5 of 35 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

SafeRI: Recognition and Intervention for Token-Level Safety Intervention in Large Vision Language Models

Existing safety alignment methods for vision-language models usually modify the model behavior globally: once the safety parameters are trained or loaded, they participate in both unsafe and already-safe generations. This always-on intervention can unnecessarily perturb the model's original reasoning path and degrade g...

Cao-Yuan Ma, Tian Gu, Wen-Pu Liu et al. · 0 citations
Preprint Sep 2026

WorldSculpt: Generating Compositional Worlds from Grounded Videos

We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics....

Mu-Yao Niu, Ji-Xuan He, Rui-Han Yu et al. · 0 citations
Preprint Aug 2026

Optical Flow from Photons

This work proposes QuantaFlow, the first method for dense optical flow directly from SPAD streams, which embeds SPAD representation construction into iterative flow refinement and constructs multi-scale representations containing intensity and structural cues.

Wendi Liu, Weichao Zeng, Wei-Hang Ran et al. · 0 citations
Jul 2026

Surprise Forcing: What to Remember, When to Skip in Long Video Generation

This work introduces Surprise Forcing, a training-free framework that treats both limitations as online resource-allocation problems and improves long-horizon consistency and visual quality while retaining real-time streaming throughput.

Shuwei Shi, Zhen Li, Mu-Yao Niu et al. · 1 citation
Preprint Aug 2026

Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation

This work proposes an adapter-based framework that incorporates event-derived cues into a pre-trained image-to-video diffusion model with minimal architectural changes and consistently outperforms existing state-of-the-art approaches.

Gui-Xu Lin, Yu-Yang Yu, Xiang Ji et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.