In this work, algorithms for classifying ground vehicles using acoustic and seismic measurements are presented. Classification based on these modalities enables identification of vehicle types through sound and ground vibration characteristics and is critical for enhancing Warfighter situational awareness in hostile environments. Artificial intelligence and machine learning techniques are used to develop a robust framework that integrates acoustic and seismic data to distinguish between military and nonmilitary vehicles. Performance is compared across traditional single-input machine learning and deep learning approaches and a multiple-input vertical federated learning framework, with emphasis on vertical federated averaging. Results show that acoustic data alone consistently yields higher classification accuracy than seismic data alone for binary military versus nonmilitary classification. Integrating acoustic and seismic data produces the highest accuracy of 95% using Extreme Gradient Boosting (XGBoost). When the military class is expanded to include wheeled vehicles (amphibious assault vehicles) and tracked vehicles (Dragon Wagons), increasing label dimensionality, XGBoost accuracy drops to 73%. In contrast, convolutional neural network (CNN)-based models achieve accuracies above 85%. A federated averaging approach using CNNs as local classifiers attains 90% accuracy. These findings indicate that multimodal vehicle classification with higher class dimensionality requires more complex learning models.
Abdoulaye Barry, Max Denis· Journal of the Acoustical So...· 0 citations
Urban environments pose significant challenges for real-time target localization and tracking due to reverberation, occlusions, and frequent GPS degradation. Visual-only systems often suffer from delayed target acquisition when line-of-sight is obstructed, while acoustic sensing provides passive, low-latency directional cues but lacks spatial resolution for persistent tracking. This work presents a real-time, edge-deployable multimodal framework that integrates acoustic sound source localization (SSL) with visual object detection and tracking to close this initialization gap. Acoustic cues are employed to actively steer a servo-mounted camera, reducing the visual search space and accelerating time-to-acquisition. Visual detections are performed using YOLOv8, while temporal consistency is maintained via BoT-SORT with camera motion compensation. An extended Kalman Filter (EKF) fuses acoustic DoA estimates and visual bounding box measurements to enable robust tracking under nonlinear motion and temporary visual occlusions.
Chidera Igwebuike, Akemi Vasquez, Wagdy H. Mahmoud et al.· Journal of the Acoustical So...· 0 citations