Do Emotional Cues Matter? Exploring Support Strategy Selection in LLM-Based Supportive Conversations
Abstract
Multimodal human–AI systems increasingly accept both text and speech, yet speech carries paralinguistic emotional cues that text does not. While prior work has evaluated the quality of large language model (LLM) responses, little is known about how vocal emotional cues reshape the support strategies an LLM selects and utilises, or whether this effect depends on emotional context. Grounding our analysis in interpersonal capitalisation theory and Hill’s helping-skills framework, we generated LLM responses to 1,132 user utterances drawn from positive (EmpatheticDialogues) and negative (ESConv) conversations, each delivered as matched text and emotion-rendered audio. A fine-tuned classifier (SSP-BERT) labelled the support strategy of every response. Vocal cues markedly shifted strategy selection in a context-dependent way. We argue that emotion recognition alone is insufficient; context-sensitive strategy adaptation is required for safe and effective multimodal support. An LLM-as-a-Judge approach was applied to evaluate human and AI responses across several criteria. However, preliminary results revealed a potential bias, with human responses receiving substantially lower ratings than LLM-generated responses across all four criteria. We are currently collecting human evaluations of the same responses to address this potential bias.