Don’t Just Fill the Silence: Exploring Conversational Filler Strategies in Embodied Virtual Agents
Abstract
Conversational fillers, such as short vocalizations (e.g., “uh,” “um”) or brief phrases that manage conversational flow, are increasingly used in embodied virtual agents. Yet, their effects on users’ perceptions remain insufficiently understood. Thus, we investigated how different conversational filler strategies shape users’ perceptions and experiences during interaction with an embodied virtual agent. We implemented a large language model (LLM)-driven conversational system that selects fillers from predefined sets and integrates them with generated responses via a prompt-engineered mechanism that maintains continuity between the filler and the response. In a within-group study ( \(N=24\) ), participants interacted with the agent across four conditions: filler-free (baseline), short, general, and context-specific fillers. Results showed that all filler types significantly improved perceived response time and reduced estimated response time relative to the filler-free condition. However, context-specific fillers provided additional benefits, including higher perceived filler quality, rapport, likability, willingness for future interaction, and lower uncanny valley ratings. Additional analyses showed that higher perceived filler quality and more favorable perceived response time were correlated with more positive perceptions of the agent and the interaction experience. These findings suggest that conversational fillers serve two roles: as temporal signals that reduce perceived delay and as context-sensitive cues that shape social perception, with meaningful benefits emerging only when fillers are contextually aligned.