The evaluation of Conversational Recommender Systems necessitates robust protocols to measure utility and user satisfaction. While human-in-the-loop testing remains the gold standard, scalability and reproducibility constraints have driven the field toward User Simulators. However, current simulation paradigms predominantly utilize rigid templates or closed-source Large Language Models that exhibit idealized behaviors. These approaches fail to capture user ambiguity, resulting in benchmarks that overestimate system proficiency by assuming crystallized user intent. To address this limitation, we introduce a family of open-weight user simulation models capable of generalizing across diverse e-commerce domains. Leveraging Teacher-Student distillation, we operationalize three distinct behavioral stereotypes:
Direct
,
Vague-Proactive
, and
Vague-Reactive
. Our evaluation of state-of-the-art Agentic Generative Conversational Recommender Systems reveals a critical
Robustness Gap
: while agents perform proficiently with decisive users, performance collapses when facing passivity and ambiguity. These findings underscore the necessity of our scalable framework for rigorously stress-testing the next generation of conversational agents against realistic, non-cooperative user behaviors.
Alessandro Petruzzelli, Alessandro Francesco Maria Martina, C. Musto et al.· Information Systems Frontier...· 0 citations
Large Language Models are increasingly used for embedding extraction. In fact, there are many approaches that try to optimize the embedding representations that these models can learn, exploiting the knowledge gained from extensive pre-training and large parameter counts. However, most works currently focus on the English language and textual input only, reflecting the trend of current Large Language Model training corpora. Recently, several Large Vision-Language Models, which are Large Language Models capable of processing multimodal signals in input, have been released. Yet, their training procedure still remains predominantly based on English data. This limitation also affects the evaluation step, with embedding benchmarks that provide limited coverage for low-resource languages. To address these challenges, we adapt a Large Vision-Language Embedding Model trained on English multimodal tasks to support multilingual inputs. Furthermore, we introduce a new benchmark to evaluate the multilingual and multimodal capabilities of embedding models.
The proposed REKALM, a comprehensive integration framework for enhancing LLM-based recommenders through knowledge integration, demonstrates that augmenting LLMs with lexicalized, domain-specific knowledge is an effective system-level strategy for advancing the next generation of recommender systems.
Alessandro Petruzzelli, C. Musto, Marco de Gemmis et al.· ACM Transactions on Informat...· 0 citations