Dialogue-Based Streaming Audio-Visual Target Speaker Extraction with Predictive Dialogue Information
In face-to-face, real-time communication, a talking agent must track the target speaker through natural pauses, turn-taking, and backchannels, often amid background cross-talk. Most target speaker extraction (TSE) studies, however, rely on simulated mixtures with full or sparse overlap and ignore the turn-taking of rea...