DEEP NEURAL FEATURE LEARNING FOR PARKINSON’S DISEASE DIAGNOSIS FROM TABULAR SPEECH DESCRIPTORS
Abstract
Parkinson’s disease (PD) is a progressive neurological disorder that can affect speech production through abnormalities in phonation, articulation, and vocal stability. This study proposes a machine learning framework for PD diagnosis using tabular speech descriptors extracted from a benchmark speech dataset. The framework includes data cleaning, missing-value inspection, feature standardization, and leakage-safe patient-wise validation to ensure reliable performance estimation. Three machine learning classifiers, namely XGBoost, LightGBM, and Support Vector Machine with Radial Basis Function (SVM-RBF), were applied and compared. Experimental evaluation was conducted using patient-wise cross-validation and assessed through multiple performance metrics, including accuracy, precision, recall, F1-score, ROC-AUC, balanced accuracy, and Matthews correlation coefficient. The results showed that LightGBM achieved the best overall performance, with an accuracy of 84.39%, precision of 85.50%, recall of 95.40%, F1-score of 90.14%, and ROC-AUC of 87.69%, outperforming XGBoost and SVM-RBF. Extended evaluation further confirmed that LightGBM provided the most balanced classification behavior, while SVM-RBF achieved the lowest false-negative rate but lower specificity. The findings indicate that boosting-based models are highly effective for detecting Parkinson’s disease from structured speech features and that leakage-safe validation is essential to obtain credible diagnostic results. The proposed framework offers a robust and reproducible approach to speech-based PD diagnosis and provides a practical foundation for future research on interpretable, externally validated clinical decision-support systems.