Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark
Conversational context provides semantic and acoustic cues across turns for automatic speech recognition (ASR), but relying on historical transcripts can propagate recognition errors and discard pronunciation and speaker information. We present a multimodal conversational-context framework for LLM-based ASR that integr...