Skip to content
Open access

Hear the Sweet Spot: Tennis Impact Localization via Single-Channel Audio

Jul 2026 · Applied Sciences · Vol 16, pp. 7340 · 0 citations · 30 references

TL;DR

A sound-sensor-based multi-task framework for racket impact localization using acoustic signals, combining radial region classification with continuous position regression, and a multi-expert convolutional neural network architecture with multi-scale feature extraction and task-specific optimization is proposed.

Abstract

Identifying the impact location (“sweet spot”) on a tennis racket is crucial for performance evaluation in tennis training. However, existing approaches typically rely on expensive vision-based systems or specialized sensors, limiting their applicability in real-world scenarios. We propose a sound-sensor-based multi-task framework for racket impact localization using acoustic signals, combining radial region classification with continuous position regression. To effectively model complex acoustic patterns, we design a multi-expert convolutional neural network (CNN) architecture with multi-scale feature extraction and task-specific optimization. Each expert branch operates at a different temporal receptive field and is trained with tailored loss functions, enabling complementary learning of global patterns, class imbalance characteristics, and hard samples. The shared backbone jointly supports both classification and regression tasks, allowing the model to learn more informative and structured representations. Experimental results demonstrate that the proposed framework consistently outperforms conventional methods in radial region classification while achieving accurate impact position estimation. Furthermore, additive noise augmentation significantly improves robustness, enabling stable performance under noisy and practical sensing conditions.

Read PDF

Similar papers

Sep 2026

Bridging the Synthetic-to-Real Gap in Impulsive Sound Detection Using Audio Transfer Learning

A CNN achieving 98.4% F1 on synthetic benchmark spectrogram data collapses to 20.0% F1 on real-world audio, revealing a severe synthetic-to-real domain gap in impulsive sound detection. This paper provides one of the first quantitative studies of this phenomenon and demonstrates that representation choice dominates cla...

Charlie Holden, Mathew Ridgely, Jayanth Bhansali et al. · 0 citations
#machine learning Preprint Sep 2026

Cross-Modal Knowledge Distillation for Acoustic Pedestrian Detection

Trust-Filtered Distillation (TFD) is introduced, which selectively suppresses teacher supervision on pedestrian samples, and its logit formulation as conditional label smoothing under a shared temperature is interpreted as conditional label smoothing under a shared temperature.

Yonghyun Kim, Chaeyeon Han, Sancho Gatungay et al. · 0 citations
Preprint Sep 2026

Disentangled Global-Local Feature Learning with E-Branchformer for Audio Deepfake Detection

The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervis...

Phuong Dat, Học Thủ, T. Nguyễn et al. · 0 citations
Open access Aug 2026

AV-DeepFake-Net: Attention-Guided and Uncertainty-Aware Network for Audiovisual DeepFake Detection

Detecting audio-visual DeepFake (AVDeepFake) is becoming increasingly important as synthetic media tools become widely accessible and spread across consumer devices. In this study, we present an Detecting audio-visual DeepFakes (AV-DeepFakes) has become increasingly critical with the rapid proliferation of accessible s...

Nasir Saleem, Ahmad Ali, Zhuo-Qi Zeng et al. · 0 citations
Open access Sep 2026

LGF-Net: a spatial-frequency deepfake detection network with learnable gabor filters and gated cross-modal Fusion

LGF-Net is proposed, a unified framework that jointly models spatial semantics and adaptive spectral cues for deepfake detection and achieves competitive intra-dataset and cross-dataset performance compared with several state-of-the-art deepfake detection methods.

Mukesh Pandey, Deepika Koundal, Sanjeev Kumar · 0 citations
Open access Sep 2026

ActNet: focus-aware multi-scale CNN for human activity recognition from images

This article proposes ActNet, a novel deep convolutional neural network architecture that combines multi-scale feature learning with a focus-aware attention mechanism to address the problem of recognizing human actions from still images.

Şafak Kılıç · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.