Skip to content
Open access

PTT-GGF Network: Robust Voice Activity Detection via Multimodal Gated Fusion in Complex Cockpit Environments

2026 · IEEE Access · Vol 14, pp. 133897-133910 · 0 citations · 35 references

TL;DR

This paper proposes a Push-to-Talk Guided Gated Fusion (PTT-GGF) audio-visual multimodal architecture that significantly outperforms mainstream audio-only and cascaded schemes, providing a highly reliable and low-latency detection paradigm for smart cockpit monitoring systems.

Abstract

Voice activity detection is a fundamental and critical component for automatic speech recognition and understanding applications in air traffic control. However, high-decibel, non-stationary noise in civil aviation cockpits severely degrades the signal-to-noise ratio, causing traditional single-modal voice activity detection algorithms to suffer from high false alarm rates and significant performance degradation. To address this critical issue and manage complex communication intentions, this paper proposes a Push-to-Talk Guided Gated Fusion (PTT-GGF) audio-visual multimodal architecture. First, we construct a temporally aligned Cockpit-AV multimodal dataset using a flight simulator, where a high-fidelity virtual Push-to-Talk logical prior is rigorously generated based on aviation Standard Operating Procedures to handle the highly sparse distribution of speech events during long-haul flights. Second, we design a confidence-aware cross-modal dynamic soft-gating strategy. Rather than simple feature concatenation, this mechanism jointly evaluates acoustic, visual, and Push-to-Talk logical states in real-time, adaptively shifting the decision focus to noise-immune visual features under high acoustic degradation. Experimental results demonstrate that the proposed PTT-GGF effectively mitigates the majority-class bias inherent in highly imbalanced data, maintaining a stable F1-Score of 95.09% under a 1:47 positive-to-negative sample ratio. Furthermore, stress tests reveal its robust performance under extreme conditions, including a -15dB signal-to-noise ratio, dynamic temporal misalignments, and transmission frame losses. With a core fusion module real-time factor of 0.000159, our method significantly outperforms mainstream audio-only and cascaded schemes, providing a highly reliable and low-latency detection paradigm for smart cockpit monitoring systems.

Read PDF

Similar papers

Open access 2026

Research on an Enhanced Zipformer With Multi-Feature Attention for Aviation Speech Recognition

Automatic Speech Recognition (ASR) for Air Traffic Control (ATC) is challenging due to factors such as fast speech rates, significant background noise, domain-specific terminology, and limited annotated data. These factors complicate the development of models that are both highly reliable and low-latency. Although the...

Geng Yu, Jian-Xing Liang, Li-Hua Wen et al. · 0 citations
Open access Sep 2026

Speech State Analysis Based on Self-Supervised Representation Shift Under Complex Speech Interference

Safety-critical control environments, such as nuclear power plant main control rooms, involve multi-operator collaboration, continuous information exchange, and stringent reliability requirements, making operator-state monitoring important for task performance and system safety. Speech is naturally produced, non-intrus...

Fu-Rui Zhao, Zhi-Hui Xu, Xing-Wei Zhang et al. · 0 citations
Open access Sep 2026

Preprocessing and Capability–Risk Coupling in an Event-Synchronized Dual-Frontend Speech Sensor System

An intelligent speech sensor is often evaluated at recognition output, although acquisition and preprocessing also determine what reaches control. We examine this propagation in an event-synchronized dual-frontend system. A conditional cascade separates artifact availability, evidence formation, semantic conversion, an...

Han Chen, Zhan-Yu Zhu · 0 citations
#artificial intelligence Preprint Sep 2026

What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection

The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes....

Jia-Jun Xu, Meng-Lu Li, Xiao-Ping Zhang · 0 citations
#natural language process... Preprint Sep 2026

FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech

Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and addition...

Puneet Mathur, Dinesh Manocha · 0 citations
2026

Low-Rate Voice Activity Detector Over Wireless Acoustic Sensor Networks

Voice Activity Detection (VAD) is a fundamental component that supports a wide range of audio/speech processing applications. Numerous studies have addressed VAD using a single microphone or a compact microphone array, yet their performance remains limited in distant scenarios. In this paper, we develop a novel data-dr...

De Hu, Shao-Jie Li, Qing-Ying Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.