Simulating Diverse User Behavioral Stereotypes for Evaluating Agentic Conversational Recommenders
Abstract
Agentic Conversational Recommender Systems (ACRSs) are designed to recommend through multi-turn dialogue with users whose needs are not fully formed at the outset. However, their evaluation almost exclusively relies on user simulators that instantiate users with clear, pre-formed needs, reducing the interaction to a retrieval over attributes disclosed in the initial turns of the conversation. This covers only a narrow slice of the behaviors real users exhibit, and assessing the robustness and reliability of these systems requires simulated users that span a wider range. To this end, we introduce a stereotype-conditioned, open-weight user simulator that spans three behavioral stereotypes: Direct, Vague-Proactive, and Vague-Reactive. Benchmarking four state-of-the-art ACRSs across four e-commerce domains with our simulator, three findings emerge. First, under certain stereotypes, the user stops contributing new information about the target as turns accumulate. At the same time, the agent continues to act, a regime previously unobserved, which we name Unproductive Stagnation and formalize via Preference Coverage. Second, a systematic Robustness Gap emerges: as the simulated user shifts from decisive to passive, accuracy collapses while conversations grow longer. Third, accuracy degrades more sharply than Preference Coverage does, decomposing the gap into two separable capabilities current ACRSs lack: elicitation and retrieval, which current evaluation entangles in a single score. Our simulator makes this distinction reportable and gives the field a controllable axis along which elicitation and retrieval can be measured and compared.