A pretrained text-to audiovisual generation model is adapted through source-conditioned feature modulation to jointly learn object addition and removal, and quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition.
Abstract
Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound from overlapping sources. To address this gap, we introduce \textit{AVIOBench}, a dataset comprising 37.9 hours of paired audiovisual examples spanning 1{,}878 target-object names. AVIOBench links the visual presence and acoustic contribution of each target object through a shared identity and visual mask. Our automated pipeline uses visual grounding and cross-modal consistency to select target-sound removal candidates, then jointly refines the audiovisual pairs to improve perceptual quality and cross-modal consistency. Building on this dataset, we propose \textit{AVIO}, which adapts a pretrained text-to audiovisual generation model through source-conditioned feature modulation to jointly learn object addition and removal. A reference-frame curriculum gradually reduces reference conditioning during training, enabling one model to perform instruction-only editing with optional visual guidance. Quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition, with optional reference guidance providing appearance and placement control for addition.
Visual object removal can eliminate a target from video frames, yet its acoustic trace persists in the soundtrack, causing obvious audio-visual inconsistency. Existing video inpainting models operate solely on pixels, while audio editing models, especially for the sound removal task, are typically driven by text and th...
Xin-Yue Guo, Jian-Xuan Yang, Dai-Guo Zhou et al.· 0 citations
This study proposes sound scene-to-description (SS2D), a method that trains an audio encoder to map complex sound patterns into the semantic space of a large language model, which allows the model to generate coherent and detailed descriptions that reflect interactions among sounds and their evolution over time, moving...
Yong-Gai Zhuang, Teng Fei, Yun-Yan Du et al.· Journal of the Acoustical So...· 0 citations
This framework produces both a transformation embedding and a processed-audio embedding, and it finds that the two play complementary roles: distance-based tasks favor the former, while probe-based tasks favor the latter.
Sungho Lee, Marco A. Martínez-Ramírez, Junghyun Koo et al.· 0 citations
Audio-visual segmentation (AVS) aims to segment sounding objects at the pixel level by integrating auditory and visual cues. However, existing methods are predominantly developed under the monaural setting and primarily rely on cross-modal semantic correspondence, which limits their ability to distinguish visually simi...
Sai-Jun Wang, Guan-Feng Tang, Hong-Bo Zhao et al.· 0 citations
ST-Omni-R1 is proposed, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning, and results on three public spatial-audio benchmarks indicate that its learned spatial and motion r...
Zhi Zeng, Cheng Zhang, Ze-Sheng Yang et al.· 2 citations
OmniEchoBench, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation, and a spatially aware omni-modal model, which introduces an FOA spatial encoder alongside a pretrained semantic audio pathway.
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.