Skip to content

Reasoning Beyond Transcription: Audio Language Models on Child Stuttering Speech

Sep 2026 · 1 citation · 32 references
Computer Science

TL;DR

Results show that while ALMs extract high-level meaning from stuttered speech, reasoning degrades significantly with increased usage, and instruction-guided models are instruction-guided to focus on the child, preserve clinically relevant disfluencies, and avoid adult-speech leakage.

Abstract

Child speech differs from adult speech in acoustics, prosody, and linguistic structures. Speech disfluencies (such as repetitions) further challenge automatic understanding. While Audio Language Models (ALMs) show strong semantic reasoning from speech audio, their ability to reason about disfluent child speech in mixed-speaker settings remains unexplored. We investigate this through two tasks: child-focused semantic summarization and speech entailment. Experiments use recordings of children who stutter in mixed speaker interviews without explicit speaker separation. Models are instruction-guided to focus on the child, preserve clinically relevant disfluencies, and avoid adult-speech leakage. Evaluation combines LLM-based judges and reference-based metrics, anchored by transcript-oracle baselines to isolate errors. Results show that while ALMs extract high-level meaning from stuttered speech, reasoning degrades significantly with increased

View source

Similar papers

Preprint Sep 2026

Speech Language Models for Full-Meeting Speaker Diarization: Capabilities and Limitations

This work investigates ESPnet-SpeechLM as a token-based backbone for generating SD hypotheses, formulating SD as autoregressive generation of structured tokens conditioned on acoustic input and shows that outputs generated by SpeechLMs encode useful temporal SD structure, but full-meeting SD remains limited by recordin...

Jia-Lu Li, Jinchuan Tian, Shinji Watanabe · 0 citations
Preprint Aug 2026

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct"naturalness"into a...

Oluwanifemi Bamgbose, Simon Rosen, J. Shah et al. · 0 citations
#natural language process... Preprint Oct 2026

Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis

Structured phonological representations provide an interpretable alternative to generic speech embeddings, but existing models are largely trained on adult speech. We adapt PhonoQ-2.0 to child speech using CHILDES-Aligned data and compare three alignment-supervision conditions (Adult, Adult+Child, and Child-only) acros...

Abner Hernandez, T. A. Vergara, Andreas K. Maier et al. · 0 citations
#natural language process... Preprint Sep 2026

Quantifying the Generation Modality Gap in Speech-Text Language Models

Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify because speech and text systems are typically evaluated with different metrics and trained on different data. We study the speech-text modality gap in a family of spoke...

Ju-Chieh Chou, Jia-Wei Zhou, Karen Livescu · 0 citations
Open access Sep 2026

Assessing the use of automatic speech recognition and large language models for individuals with language impairments

This work systematically evaluates edge-oriented ASR-LLM pipelines for individuals with language impairments using comparison studies and ablation experiments across aphasia, child language impairment, and dementia datasets to identify transcript errors, repetition, noise, and input length as factors affecting system p...

Ge-Lei Xu, Hao-Tong Yu, Li-Xuan Wei et al. · 0 citations
Open access Aug 2026

Model Development and Validation for Repetition Severity Assessment in Stuttered Speech Using Clinical Speech Datasets

A computational model that grades repetition severity from clinical speech recordings, evaluated on 480 audio samples from 60 adult speakers with persistent developmental stuttering, supports its use as a clinical decision-support tool.

J. N. Pooja, H. Y. Vani, R. P et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.