Existing safety alignment methods for vision-language models usually modify the model behavior globally: once the safety parameters are trained or loaded, they participate in both unsafe and already-safe generations. This always-on intervention can unnecessarily perturb the model's original reasoning path and degrade g...
Cao-Yuan Ma, Tian Gu, Wen-Pu Liu et al.· 0 citations
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics....
Mu-Yao Niu, Ji-Xuan He, Rui-Han Yu et al.· 0 citations
This work proposes QuantaFlow, the first method for dense optical flow directly from SPAD streams, which embeds SPAD representation construction into iterative flow refinement and constructs multi-scale representations containing intensity and structural cues.
Wendi Liu, Weichao Zeng, Wei-Hang Ran et al.· 0 citations
This work introduces Surprise Forcing, a training-free framework that treats both limitations as online resource-allocation problems and improves long-horizon consistency and visual quality while retaining real-time streaming throughput.
This work proposes an adapter-based framework that incorporates event-derived cues into a pre-trained image-to-video diffusion model with minimal architectural changes and consistently outperforms existing state-of-the-art approaches.
Gui-Xu Lin, Yu-Yang Yu, Xiang Ji et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.