Jul 2026· Information Systems Frontiers· 0 citations· 23 references
TL;DR
This work introduces a family of open-weight user simulation models capable of generalizing across diverse e-commerce domains and operationalizes three distinct behavioral stereotypes, highlighting the necessity of a scalable framework for rigorously stress-testing the next generation of conversational agents against realistic, non-cooperative user behaviors.
Abstract
The evaluation of Conversational Recommender Systems necessitates robust protocols to measure utility and user satisfaction. While human-in-the-loop testing remains the gold standard, scalability and reproducibility constraints have driven the field toward User Simulators. However, current simulation paradigms predominantly utilize rigid templates or closed-source Large Language Models that exhibit idealized behaviors. These approaches fail to capture user ambiguity, resulting in benchmarks that overestimate system proficiency by assuming crystallized user intent. To address this limitation, we introduce a family of open-weight user simulation models capable of generalizing across diverse e-commerce domains. Leveraging Teacher-Student distillation, we operationalize three distinct behavioral stereotypes:
Direct
,
Vague-Proactive
, and
Vague-Reactive
. Our evaluation of state-of-the-art Agentic Generative Conversational Recommender Systems reveals a critical
Robustness Gap
: while agents perform proficiently with decisive users, performance collapses when facing passivity and ambiguity. These findings underscore the necessity of our scalable framework for rigorously stress-testing the next generation of conversational agents against realistic, non-cooperative user behaviors.
RecVerse is presented, a GUI-grounded simulation agent that perceives pages through screenshots and produces faithful multi-turn trajectories and significantly outperforms existing baselines in both behavioral fidelity and intent consistency.
Conversational recommendation for e-commerce is increasingly mediated by large language models (LLMs), yet many real-world deployments operate under a stricter requirement: recommendations must be drawn only from a merchant's fixed catalog, without web search or unsupported product claims. In this setting, the main cha...
Ju-Li Huang, Hanna Clay, Sajjad Beygi et al.· 0 citations
Recommender systems are widely used to help users navigate information overload in digital environments. While they are often portrayed as tools that enhance autonomy by optimizing choice and personalizing content, this article argues that many current recommender systems designs in fact undermine user autonomy. Draw...
Francisco Lara, J. D. del Valle· Science and Engineering Ethi...· 0 citations
Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral signals (e.g., clicks, dwell time) rather than explicit, natural language inputs. As a result, users experience a persistent discrepancy between their explicit interests and what passiv...
Zi-Yun Xu, Bo-Sen Ding, Yue Zhang et al.· Proceedings of the 20th ACM...· 1 citation
Evaluating tool-augmented LLM agents requires diverse, realistic user inputs yet most evaluation frameworks use flat role descriptions ("you are an angry customer") that produce near-identical conversations regardless of the underlying scenario. In this paper, we propose a three-tier persona vector with 23 operationali...
Rahul Khedar, Eshita, S. R. Thondapu et al.· 0 citations
This work proposes a new approach that quantifies the effectiveness of each interaction by the reduction in the assistant's uncertainty, measured via entropy over recommendations, to fine-tune the LLM, enabling strategic interaction generation.
Cedar Site Bai, Zhen-Yu Liao, Duan Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.