Skip to content

Author

Yuxuan Chen

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

A Multimodal Emotional Interaction Framework Driven by Large Language Models

Natural and empathetic human-robot interaction is essential for social robots and other AI applications, while emotional feedback is less discussed. Thus, this paper proposes a multimodal interaction system driven by large language model (LLM). The system constructs a unified emotional state vector by integrating visual (facial expressions) and auditory (speech emotion) cues. It employs DeepSeek LLM for context-aware chain-of-thought reasoning to generate contextually appropriate verbal responses, facial expressions, and head movement commands. To achieve optimized latency and fluid embodied interaction, the system adopts a layered architecture: the upper-level LLM handles semantic understanding and behavior planning, outputting structured JSON commands; the lower-level controller generates smooth motion trajectories and manages multimodal interaction flows using a finite state machine (FSM). Experimental results demonstrate the system’s ability to effectively resolve emotional ambiguities, track emotional evolution during continuous dialogue, and achieve an optimized end-to-end response latency. This validates its feasibility and engineering value in practical human-robot interaction scenarios.

Yuxuan Chen, Chen-Yi Qiu, Ning-Chuan Wang et al. · 0 citations