AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes
A pretrained text-to audiovisual generation model is adapted through source-conditioned feature modulation to jointly learn object addition and removal, and quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition.