2026· SIGDIAL Conferences· pp. 107-123· 0 citations· 35 references
Computer Science
TL;DR
The results suggest that maintaining goal alignment in role-playing agents requires explicitly managing the trade-off between conversational responsiveness and goal adherence, and reveal a critical trade-off dependent on model scale.
Abstract
As large language models (LLMs) are increasingly deployed as role-playing agents in educational and professional training simulations, their susceptibility to user-induced distraction threatens their pedagogical utility. We formalise goal-competing distraction as a controlled evaluation paradigm and introduce a simulation framework that captures both immediate reactions and multi-turn trajectories under targeted distraction, using an LLM-based user simulator. Building upon this framework, we evaluate agent behaviour across three models: Gemini-2.0-Flash, Llama-3.3-70B-Instruct, and Llama-3.1-8B-Instruct. Our findings reveal a critical trade-off dependent on model scale. While larger models tend to remain socially responsive and more frequently engage with distractor topics, the smaller model shows rigid goal adherence by resisting and rejecting distraction. Although redirection is the most common initial response, subsequent trajectories di-verge substantially. The inclusion of explicit dialogue state demonstrates model-dependent effects, acting as a stabilising anchor for smaller models but providing limited benefit for larger ones. These results suggest that maintaining goal alignment in role-playing agents requires explicitly managing the trade-off between conversational responsiveness and goal adherence.
BluePRINT is introduced, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module, and Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory.
Si-Yu Chen, Hao-Ran Wang, Xiaojian Li et al.· 0 citations
A virtual classroom of 20 student agents who interact through rule-based chats, quarrels and consultations with friends and, when stressed, may instead consult a counselor AI under one of six style prompts, specifying the agent dynamics completely and discussing the limits of an LLM as generator of state updates.
A dual-level evaluation framework to assess LLM-based agents at both the individual and collective levels is proposed, finding that while agents capture broad partisan orientations, they underestimate within-group variability and reproduce stereotypical ideological biases.
M. Al Ali, Filip Mihai Muntean, Lucia Donatelli et al.· International Conference on...· 1 citation
This work investigates whether personality-aware fine-tuning can reduce the consistency and controllability of personality-conditioned dialogue generation compared with instruction prompting alone, and indicates that fine-tuned models are not better at role-playing different personalities than their respective baseline...
Large language models (LLMs) are increasingly used for emotional support, raising concern that sustained use may draw users away from their real-world relationships. Yet existing evaluations primarily focus on the safety, empathy, or helpfulness of responses, leaving under-examined a relational question: where does the...
Shivam Shukla, Jihye Kim, Shubham Gaur et al.· 0 citations
These findings motivate a state-action controller that starts from the base model action and selectively edits it using structural readouts, without requiring a complete predicted belief state as an intermediate representation.
Atahan Dokme, Larry Heck· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.