Estimation of voice quality parameters of dysarthric speech with preserved identity using different time-frequency image representations in deep learning network
Abstract
Speech-language pathologists use voice quality metrics of raw speech to assess dysarthria. Instead of using raw speech, a deep learning based diagnostic method that extracts these metrics from time-frequency speech representations will be reliable and preserves speaker identity. In this work a regression-based deep convolutional neural network is experimented to estimate jitter, shimmer, fundamental frequency (F0), and harmonic-to-noise ratio (HNR) employing 6 different time-frequency representations: spectrogram, low-frequency spectrogram, cepstrogram, low-frequency cepstrogram, cochleagram, and Mel scalogram. The experiment is assessed on VOC-ALS dysarthria speech dataset using RMSE and $$\text {R}^{2}$$ metrics, in which low-frequency cepstrogram performs well across all vowels, with average RMSEs of 0.76% for jitter, 2.36% for shimmer, and 4.075 dB for HNR whereas for F0, the cepstrogram shows the highest precision, with an average RMSE of 21.05 Hz. On Parkinson’s speech dataset (PC-GITA), the cepstrogram yields the lowest average RMSEs for jitter (0.57%), F0 (46.65 Hz), and HNR (3.80 dB), closely followed by the low-frequency cepstrogram with average RMSE for jitter (0.58%), F0 (48.53 Hz), and HNR (4.10 dB). For shimmer, the low-frequency cepstrogram attains the lowest average RMSE (2.64%). The results show that low-frequency cepstrogram and cepstrogram representation performs the best across all vowels for both dysarthric and Parkinson’s speech.