Skip to content
Conference

A Multimodal Emotional Interaction Framework Driven by Large Language Models

Jul 2026 · 2026 IEEE/ASME International Conference on Advanced Intelligent Mechatronics (AIM) · pp. 1-7 · 0 citations · 18 references

Abstract

Natural and empathetic human-robot interaction is essential for social robots and other AI applications, while emotional feedback is less discussed. Thus, this paper proposes a multimodal interaction system driven by large language model (LLM). The system constructs a unified emotional state vector by integrating visual (facial expressions) and auditory (speech emotion) cues. It employs DeepSeek LLM for context-aware chain-of-thought reasoning to generate contextually appropriate verbal responses, facial expressions, and head movement commands. To achieve optimized latency and fluid embodied interaction, the system adopts a layered architecture: the upper-level LLM handles semantic understanding and behavior planning, outputting structured JSON commands; the lower-level controller generates smooth motion trajectories and manages multimodal interaction flows using a finite state machine (FSM). Experimental results demonstrate the system’s ability to effectively resolve emotional ambiguities, track emotional evolution during continuous dialogue, and achieve an optimized end-to-end response latency. This validates its feasibility and engineering value in practical human-robot interaction scenarios.

View source