This paper describes TalTech's submissions to the Beyond Transcription Challenge (BeTraC), which requires generating SOAP notes directly from long doctor-patient conversation recordings, without intermediate transcription. After screening open-weight speech LLMs for long-audio robustness, we adapted Voxtral Mini (lightweight track) and Voxtral Small (heavyweight track) with LoRA supervised fine-tuning followed by DAPO reinforcement learning that uses the challenge metric, Open Medical Concept F1, as its reward. Our systems ranked first in both tracks, and an independent LLM-as-a-judge evaluation showed the lowest hallucination rate among all submissions, indicating that reinforcement learning against a concept-matching metric need not compromise factual reliability. We also find that fine-tuning on text transcripts transfers well to speech input and appears to improve robustness on out-of-domain real recordings.
Overall, matched-text speech delivery should be treated as a first-class factor in Audio LLM safety evaluation by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary.
This work introduces CoSy, a novel framework for generating diverse, steerable, multi-turn conversations at scale and evaluates CoSy on conversational grounded reasoning tasks (i.e., answering questions based on contextual information), a core on-device use case.
Patrick Huber, Arash Einolghozati, Rylan Conway et al.· IEEE Games Entertainment Med...· 0 citations
Results show that supervised fine-tuning provides the largest gain, while synthetic-speech LoRA adaptation and reinforcement learning further improve robustness, while synthetic-speech LoRA adaptation and reinforcement learning further improve robustness.
Hao Wu, Rong-Qi Han, Zhen Wang et al.· 0 citations
DCASE~2026 Task~5 introduces Audio-Dependent Question Answering (ADQA), which tests whether large audio-language models answer from the audio rather than from textual priors. An Audio-Dependency Filtering (ADF) pipeline combines silent-audio probing, per-option perplexity, a large language model (LLM) commonsense check, and human review to remove items solvable from text alone. The 3000 items that pass form the ADQA-Bench evaluation set, spanning music, speech, and environmental audio. The inaugural edition draws 14 teams and 36 submissions across two tracks defined by total parameter count (up to 100B and under 10B). A Chung-Ang University ensemble of MOSS-Audio-8B-Thinking and Qwen3-Omni-30B reaches the top overall accuracy at \pct{58.33}, and a MOSS-only configuration from the same team leads the sub-10B track at \pct{57.30}. Across the 30 submissions with a comparable development score, evaluation accuracy falls by 11.91 percentage points (pp) on average (median 10.91\,pp) on the hidden evaluation split, which is designed to be harder than the development split. The most common building blocks are: the MOSS-Audio-8B-Thinking backbone (13 of 36 submissions), Low-Rank Adaptation (LoRA) fine-tuning on AudioMCQ-StrongAC, and preference or reinforcement-learning objectives -- Group Relative Policy Optimization (GRPO) in five teams, Group reward-Decoupled Normalization Policy Optimization (GDPO) in two. At test time, prompt engineering is near-universal, and majority or choice-permutation voting is common. Every system misses the same set of 233 evaluation items.
Haolin He, Renhe Sun, Zheqi Dai et al.· 0 citations
An automated pipeline for transcription and summarization of video conferences held on the Jitsi Meet platform that requires no proprietary cloud API keys and is deployable on-premise via Docker Compose, making it suitable for organizations with strict data-privacy requirements.
G. Amirkhanova, L. Bektemir, Shyrailym Adilkyzy et al.· International Conference on...· 0 citations
This work compares the capability of generative LLMs under a five-condition prompt ablation against a fine-tuned DistilBERT token classifier at detecting self-repairs and suggests that a locally deployable encoder, given sufficient in-domain annotation, is a more plausible route to clinical self-repair detection than scaling model size or prompt complexity.
R. Wu, S. Pugh, K. O'Connor et al.· medRxiv· 0 citations