More Than a Voice: The Role of Embodiment in LLM-Based Reading Tutors for Children
Children’s engagement in reading practice is strongly influenced by social and affective factors, yet many digital reading tools lack the ability to support meaningful social interaction. In this work, we investigate how embodiment in generative AI tutors is associated with children’s social, perceptual, and emotional responses during reading activities. We developed two versions of a reading tutor powered by a Large Language Model capable of dynamically generating reading content: (1) an embodied agent with a visual avatar and (2) a non-embodied agent. In a within-subject study with children, we evaluated the impact of embodiment using eye-tracking (visual attention), self-reported measures (social presence and emotional valence), and automated facial emotion recognition. Results showed that the embodied agent was associated with higher levels of visual attention to the task and increased perceived social presence. While self-reported emotional valence did not differ significantly between conditions, the embodied agent elicited a higher proportion of positive emotional expressions during the interaction, with a moderate effect size. These findings suggest that visual embodiment may influence how children attend to and perceive AI-based tutors, supporting a more socially oriented interaction. By complementing functional feedback with responsive visual cues, embodied agents may promote engagement and observable affective responses during learning activities, highlighting their potential for educational human–AI interaction.