Skip to content
Preprint

PLACE: Positional Latent Adaptation via Conditioned Embeddings for Binaural Audio Generation

Sep 2026 · 0 citations · 27 references
Computer Science Engineering

Abstract

We present PLACE, a method that extends the pretrained any-to-audio model AudioX for binaural generation from arbitrary combinations of text, video, and optional audio prompts. PLACE augments video conditioning with Perception Encoder Core features, aligns text and video representations to derive spatial cues, and applies a conditioning-dependent low-rank transformation to the generated latent. The adapter is supervised via decoded-audio interaural level and time difference objectives. Trained on MRSAudio, PLACE improves most metrics over ViSAGe on FAIR-Play and achieves an improved SpatialCLAP score over SpatialSonic on the BEWO-1M Single Static test split. Listener evaluations favor PLACE for video-to-audio and out-of-distribution text-to-audio generation, demonstrating flexible multimodal control and improved spatial consistency.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.