Visual object removal can eliminate a target from video frames, yet its acoustic trace persists in the soundtrack, causing obvious audio-visual inconsistency. Existing video inpainting models operate solely on pixels, while audio editing models, especially for the sound removal task, are typically driven by text and th...
Xin-Yue Guo, Jian-Xuan Yang, Dai-Guo Zhou et al.· 0 citations
DuS-DiFuse is proposed, a robust dual-stream latent diffusion framework composed of a diffusion fusion unit and a generative modulation unit that achieves leading fusion performance, exhibits strong robustness to heterogeneous degradations, generalizes well across fusion tasks, and supports effective controllable gener...
Lei Cao, Hao Zhang, Peng Zhang et al.· IEEE Transactions on Pattern...· 0 citations
GRNEdit, a lightweight two-stage framework for instruction-based general video editing that outperforms multiple 14B open-source editors, while the 8B model performs on par with leading open-source editors.
TetherMem is introduced, a training-free, query-aware spatiotemporal memory router for frozen video generators that separates subject and scene queries and modulates historical access with region- and age-conditioned priors: subject queries retain identity-bearing history, while scene queries reduce reliance on subject...
Chen Li, Peng Zhang, Han-Yu Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.