A sound-sensor-based multi-task framework for racket impact localization using acoustic signals, combining radial region classification with continuous position regression, and a multi-expert convolutional neural network architecture with multi-scale feature extraction and task-specific optimization is proposed.
Abstract
Identifying the impact location (“sweet spot”) on a tennis racket is crucial for performance evaluation in tennis training. However, existing approaches typically rely on expensive vision-based systems or specialized sensors, limiting their applicability in real-world scenarios. We propose a sound-sensor-based multi-task framework for racket impact localization using acoustic signals, combining radial region classification with continuous position regression. To effectively model complex acoustic patterns, we design a multi-expert convolutional neural network (CNN) architecture with multi-scale feature extraction and task-specific optimization. Each expert branch operates at a different temporal receptive field and is trained with tailored loss functions, enabling complementary learning of global patterns, class imbalance characteristics, and hard samples. The shared backbone jointly supports both classification and regression tasks, allowing the model to learn more informative and structured representations. Experimental results demonstrate that the proposed framework consistently outperforms conventional methods in radial region classification while achieving accurate impact position estimation. Furthermore, additive noise augmentation significantly improves robustness, enabling stable performance under noisy and practical sensing conditions.
A CNN achieving 98.4% F1 on synthetic benchmark spectrogram data collapses to 20.0% F1 on real-world audio, revealing a severe synthetic-to-real domain gap in impulsive sound detection. This paper provides one of the first quantitative studies of this phenomenon and demonstrates that representation choice dominates cla...
Charlie Holden, Mathew Ridgely, Jayanth Bhansali et al.· International Symposium on N...· 0 citations
Trust-Filtered Distillation (TFD) is introduced, which selectively suppresses teacher supervision on pedestrian samples, and its logit formulation as conditional label smoothing under a shared temperature is interpreted as conditional label smoothing under a shared temperature.
Yonghyun Kim, Chaeyeon Han, Sancho Gatungay et al.· 0 citations
The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervis...
Phuong Dat, Học Thủ, T. Nguyễn et al.· 0 citations
Detecting audio-visual DeepFake (AVDeepFake) is becoming increasingly important as synthetic media tools become widely accessible and spread across consumer devices. In this study, we present an Detecting audio-visual DeepFakes (AV-DeepFakes) has become increasingly critical with the rapid proliferation of accessible s...
Nasir Saleem, Ahmad Ali, Zhuo-Qi Zeng et al.· International Journal of Int...· 0 citations
LGF-Net is proposed, a unified framework that jointly models spatial semantics and adaptive spectral cues for deepfake detection and achieves competitive intra-dataset and cross-dataset performance compared with several state-of-the-art deepfake detection methods.
This article proposes ActNet, a novel deep convolutional neural network architecture that combines multi-scale feature learning with a focus-aware attention mechanism to address the problem of recognizing human actions from still images.
Şafak Kılıç· PeerJ Computer Science· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.