Aug 2026· Message Understanding Conference· pp. 617-622· 0 citations· 22 references
Computer Science
TL;DR
A ready-to-deploy intent-aware system in which a social robot conveys active listening through non-verbal backchannels grounded in interactional intents to enable active listening for robots.
Abstract
In conversational Human-Robot Interaction, robots typically remain silent during user speech and reply only after a pause, making interaction feel unnatural. In contrast, humans signal that they listen through active behavior. To overcome this, we present a system in which a social robot conveys active listening through non-verbal backchannels grounded in interactional intents. The system combines two ideas: (i) a dual-stage framework separating the user’s communicative intent (Speaker Intent) from the robot’s interactional stance (Listener Intent), mapping the latter to non-verbal reactions; and (ii) a parallel pipeline whose chunk-level branch generates non-verbal feedback during speech while a turn-level branch produces the verbal reply at turn end. We deployed this system on a robot and conducted a usability study (N = 6) in a hotel-negotiation task. We found that participants considered the system usable and could interpret gestures. We contribute a ready-to-deploy intent-aware system to enable active listening for robots.
A mixed-reality conversational system that integrates voice interaction, LLM-driven dialogue, and prosody-based affect-aware adaptation with a modular client–server architecture is presented, enabling emotion-related cues inferred from vocal prosody to be incorporated without interrupting conversational flow.
Andrea Antonio Cantone, Matteo Ercolino, M. Sebillo et al.· Journal of Ambient Intellige...· 0 citations
Preliminary results suggest that explicitly modeling both speaker emotional dynamics and listener affective state can improve embodied empathetic interaction.
Z. Pang, C. Kennington, Tatsuya Kawahara· 0 citations
% !TEX root = ../main.tex Real-time embodied companion interaction requires a robot to infer user intent from streaming speech, generate timely responses, and execute expressive, interruptible motions. Existing systems typically decouple dialogue orchestration from gesture synthesis, relying on offline motion generatio...
Li-Jian Lin, Ye Zhu, Fan Zhang et al.· 0 citations
Socially competent humanoid robots must communicate affect and intent through gesture as well as speech, yet open-ended interaction must become motion that is both expressive and executable on a specific body. This demands semantic flexibility for contextual social intent while preserving deterministic, embodiment-awar...
Dao-Jie Peng, Bing-Tao Wang, Fu-Long Ma et al.· 0 citations
Turn-taking prediction is especially relevant for social robots that act as mediators in human-human interaction, where the expected action is often not to speak, but to orient, wait, avoid interruption, or prepare a balanced intervention. This paper presents Multimodal Voice Activity Projection (MM-VAP) as a human sta...
Antonio Cano, Guillermo Pérez, Luis Merino et al.· 0 citations