Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representation of where sound sources are located a...
Duo-Wen Chen, Jin-Jin He, K. Gouthaman et al.· 0 citations
Across five 5.1-channel test sets, FGAtt achieves the strongest over- all performance, demonstrating the effectiveness of feature-level adaptation for reusing pretrained perceptual knowledge for multichannel spatial audio.
K. Gouthaman, S. Gehlot, Vishnu Raj et al.· 0 citations
VIBE is introduced, a novel text-and-video-to-music (T+V2M) generation model that leverages a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints and soft, subjec...