Aug 2026· 2026 12th International Conference on Big Data and Information Analytics (BigDIA)· pp. 327-334· 0 citations· 20 references
Abstract
In modern operational environments, rapid and hands-free target localization is crucial for situational awareness. However, traditional plotting systems rely on cumbersome manual interactions, and conventional multimodal algorithms degrade significantly under extreme background noise and constrained communication links. To address these challenges, we propose a novel edge-cloud collaborative Speech-to-Plot (STP) framework. The proposed system integrates a domain-adapted Automatic Speech Recognition (ASR) module—fine-tuned via a noise-injected curriculum—with a zero-shot visual grounding model to translate natural voice commands into precise spatial bounding boxes. Evaluations on a custom domain-specific dataset demonstrate that our framework exhibits graceful degradation rather than severe degradation under extreme acoustic interference, maintaining robust target semantic extraction even at 0 dB Signal-to-Noise Ratio. This reliable acoustic front-end helps prevent cascading errors in downstream cross-modal attention mechanisms, enabling accurate visual target localization. Furthermore, stress testing under simulated narrowband communication networks validates the practical engineering viability of our decoupled architecture. By offloading heavy multimodal inference to the cloud, the system mitigates computational congestion, supporting operational resilience despite the inevitable physical bandwidth limitations of field deployments.
This paper proposes a Push-to-Talk Guided Gated Fusion (PTT-GGF) audio-visual multimodal architecture that significantly outperforms mainstream audio-only and cascaded schemes, providing a highly reliable and low-latency detection paradigm for smart cockpit monitoring systems.
Wei-Jun Pan, Qi-Xiang Wang, Yi-Di Wang et al.· IEEE Access· 0 citations
DAVE is presented, a decoupled audio-visual enhancement framework for real-world speech separation that applies scene routing, GAN-based denoising, and loudness normalization only within the no-reference partition, guaranteeing non-degradation of reference-based metrics.
Environmental noise obscures the content cues needed for whisper-to-normal conversion. We propose WhisperVC-AV, which restores acoustic content features while leaving the original WhisperVC conversion module unchanged. Its context-guided restoration module combines temporal acoustic context with synchronized lip featur...
Replay attacks are the most accessible threat to voice-controlled systems, and the acoustic cues that expose them are strongly modulated by the environment in which the attack is mounted. A detector deployed in the field therefore has to absorb new acoustic conditions over time, ideally without revisiting past recordin...
Spatial audio is a key component of immersive $360^\circ$ media, yet high-quality spatial capture remains limited in real-world speech-dominant scenes. We study visually guided First-Order Ambisonics (FOA) speech spatialization in the wild: given aligned $360^\circ$ video and an omnidirectional audio track, we recover...
Qing-Yu Luo, Peng Zhang, Wen-Wu Wang et al.· 0 citations
Video-to-speech synthesis aims to generate natural-sounding speech from silent talking-face videos while ensuring phonetic accuracy. A fundamental challenge in this task is the inherent one-to-many mapping problem, where visual dynamics often lack sufficient information to uniquely determine the corresponding utterance...
Gun-Woo Lee, Yoori Oh, Yoseob Han· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.