Skip to content
Conference

WhisperEngineClassifier: A Transfer Learning Framework for Audio-Based Engine State Classification

Aug 2026 · 2026 6th International Conference on Emerging Smart Technologies and Applications (eSmarTA) · pp. 1-8 · 0 citations · 28 references

Abstract

The process of diagnosing and monitoring automotive systems through audio-based engine state classification has become more vital for intelligent vehicle systems as real-world applications face challenges from background noise, device differences, and class distribution problems which are issues addressed in this research. The VGG-Sound engine sound database was developed through ontology-based filtering, manual annotation, preprocessing, and five-state source-disjoint splitting. Eight CNN baselines were benchmarked under a unified training setup, revealing limitations in modeling long-range temporal dependencies. The development of WhisperEngineClassifier requires the adaptation of a pre-trained Whisper speech encoder through the removal of its decoder and the addition of lightweight pooling heads which will be fine-tuned in two different phases. The proposed model achieved 85.0% accuracy and 0.850 F1-score, outperforming the best CNN baseline, DenseNet169, by 16.2% in accuracy and showing clear gains on difficult classes. The system shows its ability to operate in real-world applications through a Flask-based deployment that uses FP16 quantization and GPU acceleration to support automotive telematics and predictive maintenance.

View source