This paper proposes a Push-to-Talk Guided Gated Fusion (PTT-GGF) audio-visual multimodal architecture that significantly outperforms mainstream audio-only and cascaded schemes, providing a highly reliable and low-latency detection paradigm for smart cockpit monitoring systems.
Abstract
Voice activity detection is a fundamental and critical component for automatic speech recognition and understanding applications in air traffic control. However, high-decibel, non-stationary noise in civil aviation cockpits severely degrades the signal-to-noise ratio, causing traditional single-modal voice activity detection algorithms to suffer from high false alarm rates and significant performance degradation. To address this critical issue and manage complex communication intentions, this paper proposes a Push-to-Talk Guided Gated Fusion (PTT-GGF) audio-visual multimodal architecture. First, we construct a temporally aligned Cockpit-AV multimodal dataset using a flight simulator, where a high-fidelity virtual Push-to-Talk logical prior is rigorously generated based on aviation Standard Operating Procedures to handle the highly sparse distribution of speech events during long-haul flights. Second, we design a confidence-aware cross-modal dynamic soft-gating strategy. Rather than simple feature concatenation, this mechanism jointly evaluates acoustic, visual, and Push-to-Talk logical states in real-time, adaptively shifting the decision focus to noise-immune visual features under high acoustic degradation. Experimental results demonstrate that the proposed PTT-GGF effectively mitigates the majority-class bias inherent in highly imbalanced data, maintaining a stable F1-Score of 95.09% under a 1:47 positive-to-negative sample ratio. Furthermore, stress tests reveal its robust performance under extreme conditions, including a -15dB signal-to-noise ratio, dynamic temporal misalignments, and transmission frame losses. With a core fusion module real-time factor of 0.000159, our method significantly outperforms mainstream audio-only and cascaded schemes, providing a highly reliable and low-latency detection paradigm for smart cockpit monitoring systems.
Automatic Speech Recognition (ASR) for Air Traffic Control (ATC) is challenging due to factors such as fast speech rates, significant background noise, domain-specific terminology, and limited annotated data. These factors complicate the development of models that are both highly reliable and low-latency. Although the...
Safety-critical control environments, such as nuclear power plant main control rooms, involve multi-operator collaboration, continuous information exchange, and stringent reliability requirements, making operator-state monitoring important for task performance and system safety. Speech is naturally produced, non-intrus...
An intelligent speech sensor is often evaluated at recognition output, although acquisition and preprocessing also determine what reaches control. We examine this propagation in an event-synchronized dual-frontend system. A conditional cascade separates artifact availability, evidence formation, semantic conversion, an...
Han Chen, Zhan-Yu Zhu· Italian National Conference...· 0 citations
The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes....
Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and addition...
Voice Activity Detection (VAD) is a fundamental component that supports a wide range of audio/speech processing applications. Numerous studies have addressed VAD using a single microphone or a compact microphone array, yet their performance remains limited in distant scenarios. In this paper, we develop a novel data-dr...
De Hu, Shao-Jie Li, Qing-Ying Zhao et al.· IEEE Transactions on Audio,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.