Benchmarking Classical ML and Deep Learning Models for Parkinson's Detection: A Speech Biomarker Study with XAI
Parkinson's Disease (PD) is a neurodegenerative disorder which affects about 10 million people worldwide, and is clinically difficult to diagnose early and accurately since only subjective motor evaluation is used. The speech biomarkers are now also a valuable non-invasive diagnostic pathway, because hypokinetic dysarthria is present in almost 80–90% of patients with PD and can predate manifest motor symptoms. In this study, a hierarchical hybrid neural framework, DeepNetX2, which combines parallel components—spatial-convolutional and bidirectional recurrent—with a gated attention fusion network (gAFN) is proposed for the speech-based PD classification. The proposed framework is evaluated on UCI Parkinson's Disease Speech Features Dataset which includes 756 voice recordings and 754 acoustic features and achieves a classification accuracy of 94.74% and an AUC-ROC of 0.9607 under speaker-independent 10-fold crossvalidation with better performance than classical ML baselines such as SVM (85.09%) and Gradient Boosting (89.47%) and deep learning architectures such as CNN-LSTM (92.10%). It is designed to overcome the diagnostic black-box problem by embedding a dual-layer Explainable AI system that uses SHAP global attribution and LIME local surrogate models, ensuring two clinically significant bands (Pitch Period Entropy – PPE and Tunable Q-factor Wavelet Transform – TQWT decomposition bands) are identified. By introducing unified and standardized benchmarks across seven topologies of ML and DL and under the same pre-processing pipelines, the study fills a gap in the literature. Results show that the combination of hierarchical deep learning with the interpretable XAI wrappers provide high accuracy and clinically trustworthy diagnostic systems that can be useful in early screening of neurodegenerative diseases.