Skip to content

AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes

Sep 2026 · 0 citations · 96 references
Computer Science

TL;DR

A pretrained text-to audiovisual generation model is adapted through source-conditioned feature modulation to jointly learn object addition and removal, and quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition.

Abstract

Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound from overlapping sources. To address this gap, we introduce \textit{AVIOBench}, a dataset comprising 37.9 hours of paired audiovisual examples spanning 1{,}878 target-object names. AVIOBench links the visual presence and acoustic contribution of each target object through a shared identity and visual mask. Our automated pipeline uses visual grounding and cross-modal consistency to select target-sound removal candidates, then jointly refines the audiovisual pairs to improve perceptual quality and cross-modal consistency. Building on this dataset, we propose \textit{AVIO}, which adapts a pretrained text-to audiovisual generation model through source-conditioned feature modulation to jointly learn object addition and removal. A reference-frame curriculum gradually reduces reference conditioning during training, enabling one model to perform instruction-only editing with optional visual guidance. Quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition, with optional reference guidance providing appearance and placement control for addition.

View source

Similar papers

Preprint Sep 2026

TV-AudioRemover: Joint Text-Visual Guided Sound Removal with Multi-Task Hard-Mixture Curriculum

Visual object removal can eliminate a target from video frames, yet its acoustic trace persists in the soundtrack, causing obvious audio-visual inconsistency. Existing video inpainting models operate solely on pixels, while audio editing models, especially for the sound removal task, are typically driven by text and th...

Xin-Yue Guo, Jian-Xuan Yang, Dai-Guo Zhou et al. · 0 citations
Sep 2026

From acoustic scene to sense of place: Sound scene-to-description, a semantic translation engine for sound environment.

This study proposes sound scene-to-description (SS2D), a method that trains an audio encoder to map complex sound patterns into the semantic space of a large language model, which allows the model to generate coherent and detailed descriptions that reflect interactions among sounds and their evolution over time, moving...

Yong-Gai Zhuang, Teng Fei, Yun-Yan Du et al. · 0 citations
Preprint Aug 2026

Exploring the Design Space of Representation Learning for Audio Transformations

This framework produces both a transformation embedding and a processed-audio embedding, and it finds that the two play complementary roles: distance-based tasks favor the former, while probe-based tasks favor the latter.

Sungho Lee, Marco A. Martínez-Ramírez, Junghyun Koo et al. · 0 citations
Preprint Sep 2026

Binaural Audio-Visual Instance Segmentation

Audio-visual segmentation (AVS) aims to segment sounding objects at the pixel level by integrating auditory and visual cues. However, existing methods are predominantly developed under the monaural setting and primarily rely on cross-modal semantic correspondence, which limits their ability to distinguish visually simi...

Sai-Jun Wang, Guan-Feng Tang, Hong-Bo Zhao et al. · 0 citations
Preprint Aug 2026

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

ST-Omni-R1 is proposed, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning, and results on three public spatial-audio benchmarks indicate that its learned spatial and motion r...

Zhi Zeng, Cheng Zhang, Ze-Sheng Yang et al. · 2 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.