Exploring Multimodal Turn-Taking Cues in Face-to-Face Conversation using Voice Activity Projection
Results show that visual information improves performance over the audio-only baseline, with FAU being significantly more informative than other feature groups, and results indicate that performance on specific tasks varies depending on whether training and test data come from improvised or naturalistic conversations.