This work presents the VOR Dataset (VORD), the first benchmark dataset providing both paired edited videos and graffiti masks, and introduces VOR-MDSM, the first perception-driven VLM-based scoring model specifically designed for mask-guided VOR.
Hao-Nan Huang, Tian-Rui Qiu, Xiang-Hao Zang et al.· 0 citations
OmniMate, a unified framework for open-ended real-time interactive audio-visual avatar generation and a Multi-Reference Conditioning Module (MRCM), which leverages multiple reference images and a reference speech segment to provide persistent visual and speaker identity cues throughout long-duration streaming interacti...
Ripple is a real-time joint audio-video generation system with a cross-modal recurrent memory mechanism that combines a fixed-length sliding-window attention with modality-specific memory states that continuously summarize audio and video context.
Yan-Bo Ding, Zhi-Zhi Guo, Quan-Yue Song et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.