Skip to content
Open access

Automatic analysis of speech representations to assess psychological distress

Aug 2026 · Frontiers in Digital Health · Vol 8 · 0 citations · 13 references
Medicine

Abstract

Background Current mental health diagnostic methods are limited by subjective clinical interpretation. Automatic speech analysis is a promising technology for objective assessment. Objective To evaluate and compare different speech-based representations (acoustic, phonetic, and time-frequency) and deep learning-based embeddings for discriminating symptoms associated with psychological distress. Methods A secondary analysis of the Distress Analysis Interview Corpus (DAIC-WOZ) was conducted using recordings from 125 participants (3,069 responses). Speech representations included phonation, articulation, and prosody features extracted with DisVoice; phonetic features extracted with Phonet; time-frequency representations derived from Mexican hat wavelets; and deep embeddings extracted with the multilingual Wav2Vec 2.0 model XLSR-53. Two classification strategies were addressed at the response and participant levels using a Fully Connected Neural Network (FCNN) and a Support Vector Machine (SVM), respectively. Results Prosody at the participant level achieved the highest mean performance (F1-score 0.67 ± 0.07; accuracy 0.64 ± 0.10; AUC 0.65 ± 0.12), followed by participant-level phonation (F1-score 0.59 ± 0.16; accuracy 0.61 ± 0.14; AUC 0.65 ± 0.16). Conversely, participant-level aggregation of deep embeddings yielded lower performance (F1-score 0.48 ± 0.19; accuracy 0.55 ± 0.13; AUC 0.52 ± 0.15), failing to surpass traditional features. Response-level performance remained close to chance. Phonet and wavelet representations did not improve performance over prosody or phonation. Conclusion Participant-level analysis provided more robust and consistent discriminative patterns than response-level approaches. Prosody and phonation achieved the best performance across speech representations, while phonetic, time-frequency, and deep speech representations did not outperform the best acoustic baseline. These findings suggest that, within the evaluated experimental setting, the aggregation strategy appears to have a stronger influence on performance than increasing representational complexity.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.