Emotion-Aware Autonomous Driving: A Survey on Multimodal Emotion Recognition, LLM-Based Decision-Making, and Human-Machine Interaction
Abstract
Autonomous driving research has begun to move. Where the early emphasis fell on safety and efficiency, attention is now turning toward the people inside the car and how the ride actually feels. Most systems, though, still read the road without reading the passenger: real-time emotional states rarely enter the perception or decision pipeline, and that gap makes human-vehicle trust hard to build and the ride harder to improve. This review looks at three lines of work that bear on the problem—in-vehicle multimodal emotion recognition, driving decisions built on large language models (LLMs), and the mechanisms behind human-machine interaction and trust—and reads them against one another rather than in isolation. A consistent pattern emerges. Emotion recognition works well in controlled tests, yet its output seldom reaches the decision module. LLMs plan tasks and explain their choices, but no current system feeds passenger emotion into their reasoning. Interaction research cares about trust and explainability, only its content stays fixed and scripted, so it cannot bend to how a passenger feels in the moment. The three lines run in parallel and almost never meet. That missing coupling is the field’s central gap, and it is what keeps a vehicle from sensing, let alone answering, the subjective state of the person it carries. We then work through what integration would demand—latency, safety limits, and personalization—and set out five directions for future work: a unified perception-decision framework, dynamic explainable interaction driven by an LLM, lightweight in-vehicle deployment, personalized learning, and better datasets. Taken together, they sketch a reference for a closed loop that runs from emotion perception through decision generation to human-machine interaction.