Trust-Filtered Distillation (TFD) is introduced, which selectively suppresses teacher supervision on pedestrian samples, and its logit formulation as conditional label smoothing under a shared temperature is interpreted as conditional label smoothing under a shared temperature.
Abstract
Audio-only pedestrian detection is attractive for urban sensing but limited by weak acoustic cues. An appealing strategy is cross-modal knowledge distillation, in which a video teacher supervises the audio student during training so that the deployed model runs on audio alone. Under this task's severe class imbalance and wide video-audio modality gap, however, what such distillation contributes is unclear. We introduce Trust-Filtered Distillation (TFD), which selectively suppresses teacher supervision on pedestrian samples, and interpret its logit formulation as conditional label smoothing under a shared temperature. Across ten distillation configurations under five-fold cross-session validation on ASPED, most methods yield modest gains in macro accuracy accompanied by small changes in PR-AUC. The main effect is higher no-pedestrian accuracy at the cost of lower pedestrian recall. Adding TFD to logit distillation strengthens this trade-off without improving mean PR-AUC. These findings clarify the benefits and limitations of selective cross-modal supervision by distinguishing operating-point shifts from discrimination gains in imbalanced acoustic detection.
Knowledge distillation (KD) improves low-resource acoustic learning by enriching one-hot supervision with the softened predictive distribution of a fixed teacher network. However, a teacher trained with limited or imbalanced annotations may produce a biased distribution whose components are not uniformly reliable. Alth...
Shuang-Lin Li, Ru-Xiao Qian, Jian Liu et al.· 0 citations
Omni-modal large language models deployed in real-world environments encounter external noise that can interfere with their perception and understanding of multimodal inputs. We study their robustness in audio-visual understanding, focusing on question answering under environmental noise and competing speech. The chall...
Xing-Ming Shui, Da-Peng Chen, Bo-Wei Liu et al.· 0 citations
Emergency vehicle detection in autonomous driving is a safety-critical perception task that demands robustness under diverse and adverse real-world conditions. Existing approaches rely on a single modality, either audio or video, which leads to systematic failure when that modality is degraded: microphone-based systems...
PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis, is addressed, making it substantially faster than gradient-based TTA while requiring no additional training.
A. Shukla, R. Thakur, Aryan Das et al.· 0 citations
Acoustic and image sensors in unattended ground sensor systems often operate at different sampling rates and use different trigger mechanisms and fields of view, so their observations are asynchronous and lack instance-level correspondence. In addition, web images used to supplement scarce field annotations exhibit a s...
Yan Wang, Li-Jie Xia, Jun Pei et al.· Applied Sciences· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026