Sep 2026· Journal of Ambient Intelligence and Humanized Computing· 0 citations· 9 references
TL;DR
A mixed-reality conversational system that integrates voice interaction, LLM-driven dialogue, and prosody-based affect-aware adaptation with a modular client–server architecture is presented, enabling emotion-related cues inferred from vocal prosody to be incorporated without interrupting conversational flow.
Abstract
Conversational agents based on Large Language Models (LLMs) are increasingly explored in eXtended Reality (XR) applications that require real-time guidance during task execution. In task-oriented XR environments, conversational support must remain responsive and integrated with the ongoing activity, minimizing interruptions in user interaction. Although affect-aware adaptation has been proposed as a possible strategy to improve interaction quality, its impact in immersive task-oriented scenarios remains unclear. This paper presents a mixed-reality conversational system that integrates voice interaction, LLM-driven dialogue, and prosody-based affect-aware adaptation. The system adopts a modular client–server architecture with parallel semantic and affective processing pipelines, enabling emotion-related cues inferred from vocal prosody to be incorporated without interrupting conversational flow. The system was evaluated through a between-subjects study involving 40 participants performing a guided chemical procedure within a mixed-reality laboratory scenario. Participants interacted with either an affect-adaptive or a non-adaptive version of the conversational agent. The evaluation combined standardized questionnaires, behavioral metrics extracted from interaction logs, and qualitative feedback. Results showed that the affect-adaptive condition produced greater perceived usability scores compared to the non-adaptive condition. Participants interacting with the adaptive agent also required longer task completion times and engaged in a higher number of conversational turns. No significant differences emerged for overall workload between the two conditions. Qualitative findings further suggested that affect-aware adaptation made the interaction feel more natural and supportive.
AffAdapt is presented, a seamless interaction design framework for AI-personas, which coordinates streaming speech recognition, proactive turn-management, persona-grounded response generation, a persistent emotional state, and synchronized embodied output into a single interaction loop.
Nishanth Chidambaram, Kaustubh Paliwal, Kayla Hom et al.· 1 citation· ⚡1
A ready-to-deploy intent-aware system in which a social robot conveys active listening through non-verbal backchannels grounded in interactional intents to enable active listening for robots.
Yang Sun, Jan Leusmann, Michael A. Hedderich· Message Understanding Confer...· 0 citations
Conversational fillers, such as short vocalizations (e.g., “uh,” “um”) or brief phrases that manage conversational flow, are increasingly used in embodied virtual agents. Yet, their effects on users’ perceptions remain insufficiently understood. Thus, we investigated how different conversational filler strategies shape...
Dynamic speaking style adaptation in multi-turn multimodal interaction remains a major challenge for text-to-speech (TTS) systems. Existing context-aware TTS (CTTS) methods typically map dialogue context to speech in an end-to-end manner. Such implicit modeling makes contextual style decisions difficult to supervise, w...
Wen-Jie Tian, Kang-Xiang Xia, Jing-Bin Hu et al.· 0 citations
Voice-assistant interruptions tend to be intrusive because existing systems fail to consider the affective state, cognitive load and situational context of the user when deciding when and how to interrupt.Voice-assistant interruptions tend to be intrusive, since existing systems do not consider the affective state, cog...
Preliminary results suggest that explicitly modeling both speaker emotional dynamics and listener affective state can improve embodied empathetic interaction.
Z. Pang, C. Kennington, Tatsuya Kawahara· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.