It is argued nested model-selection protocols should be standard when probing encoder layer representations on small clinical speech datasets, and a fair comparison across representation families identifies a compact, task-aligned affect--prosody model as competitive with substantially more complex SSL, deep-learning,...
P. A. Pérez-Toro, David Gimeno-Gómez, D. Rückert et al.· 0 citations
Multimodal medical prediction often faces incomplete pairing: auxiliary modalities with complementary signal are available for only a subset of subjects (or none) and cannot be assumed at deployment. We introduce PANDA (Prototype Anchored Data Alignment), a two-stage framework that transfers auxiliary information to a...
Sheethal Bhat, Mahfuzur Rahman Chowdhury, P. A. Pérez-Toro et al.· 0 citations
Structured phonological representations provide an interpretable alternative to generic speech embeddings, but existing models are largely trained on adult speech. We adapt PhonoQ-2.0 to child speech using CHILDES-Aligned data and compare three alignment-supervision conditions (Adult, Adult+Child, and Child-only) acros...
Abner Hernandez, T. A. Vergara, Andreas K. Maier et al.· 0 citations
Speech-based Alzheimer's disease (AD) assessments increasingly rely on pretrained self-supervised learning (SSL) models that learn acoustic representations directly from raw audio, exposing the model to recording factors. We ask whether such factors are merely encoded in SSL representations or can systematically alter...
Serli Kopar, Alkis Koudounas, R. Rane et al.· 0 citations
Preliminary analyses show interpretable surface-sensitive patterns consistent with flapping-like /t/ realizations, /t/-/r/ retraction or affrication, and nasal place assimilation, indicating that phonological information from synchronized audio can be partially transferred to articulatory models.
Abner Hernandez, T. A. Vergara, Dai-Qi Liu et al.· 0 citations
A layer-wise analysis of nine SSL speech backbones using a low-capacity logistic regression probe reveals that the transferred discriminative signal lacks pathological specificity, highlighting critical limitations that must be addressed before speech-based pathology recognition models can be reliably deployed in clini...
Serli Kopar, Sam Gijsen, Abner Hernandez et al.· 0 citations
A standardization-oriented framework that turns evaluation assumptions into explicit, reproducible evidence and provides a basis for more comparable, auditable evaluation and future certification-oriented assessment of machine-learning protection functions is proposed.
Julian Oelhaf, Georg Kordowich, Paula Andrea Pérez-Toro et al.· International Journal of Ele...· 0 citations
This paper investigates how multilingual medical adaptation reshapes the internal representations of Whisper models through layer-wise encoder analysis, and shows that English medical fine-tuning produces the dominant encoder shift, whereas multilingual continuation largely preserves the adapted representation space.
Souranil Kahali, Rituparna Bose, Abner Hernandez et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.