Results show that while ALMs extract high-level meaning from stuttered speech, reasoning degrades significantly with increased usage, and instruction-guided models are instruction-guided to focus on the child, preserve clinically relevant disfluencies, and avoid adult-speech leakage.
Abstract
Child speech differs from adult speech in acoustics, prosody, and linguistic structures. Speech disfluencies (such as repetitions) further challenge automatic understanding. While Audio Language Models (ALMs) show strong semantic reasoning from speech audio, their ability to reason about disfluent child speech in mixed-speaker settings remains unexplored. We investigate this through two tasks: child-focused semantic summarization and speech entailment. Experiments use recordings of children who stutter in mixed speaker interviews without explicit speaker separation. Models are instruction-guided to focus on the child, preserve clinically relevant disfluencies, and avoid adult-speech leakage. Evaluation combines LLM-based judges and reference-based metrics, anchored by transcript-oracle baselines to isolate errors. Results show that while ALMs extract high-level meaning from stuttered speech, reasoning degrades significantly with increased
This work investigates ESPnet-SpeechLM as a token-based backbone for generating SD hypotheses, formulating SD as autoregressive generation of structured tokens conditioned on acoustic input and shows that outputs generated by SpeechLMs encode useful temporal SD structure, but full-meeting SD remains limited by recordin...
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct"naturalness"into a...
Oluwanifemi Bamgbose, Simon Rosen, J. Shah et al.· 0 citations
Structured phonological representations provide an interpretable alternative to generic speech embeddings, but existing models are largely trained on adult speech. We adapt PhonoQ-2.0 to child speech using CHILDES-Aligned data and compare three alignment-supervision conditions (Adult, Adult+Child, and Child-only) acros...
Abner Hernandez, T. A. Vergara, Andreas K. Maier et al.· 0 citations
Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify because speech and text systems are typically evaluated with different metrics and trained on different data. We study the speech-text modality gap in a family of spoke...
This work systematically evaluates edge-oriented ASR-LLM pipelines for individuals with language impairments using comparison studies and ablation experiments across aphasia, child language impairment, and dementia datasets to identify transcript errors, repetition, noise, and input length as factors affecting system p...
A computational model that grades repetition severity from clinical speech recordings, evaluated on 480 audio samples from 60 adult speakers with persistent developmental stuttering, supports its use as a clinical decision-support tool.
J. N. Pooja, H. Y. Vani, R. P et al.· International journal of com...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.