Skip to content
Open access

Breaking the Script: Do Role-Playing Agents Maintain Goal Alignment under Distraction?

2026 · SIGDIAL Conferences · pp. 107-123 · 0 citations · 35 references
Computer Science

TL;DR

The results suggest that maintaining goal alignment in role-playing agents requires explicitly managing the trade-off between conversational responsiveness and goal adherence, and reveal a critical trade-off dependent on model scale.

Abstract

As large language models (LLMs) are increasingly deployed as role-playing agents in educational and professional training simulations, their susceptibility to user-induced distraction threatens their pedagogical utility. We formalise goal-competing distraction as a controlled evaluation paradigm and introduce a simulation framework that captures both immediate reactions and multi-turn trajectories under targeted distraction, using an LLM-based user simulator. Building upon this framework, we evaluate agent behaviour across three models: Gemini-2.0-Flash, Llama-3.3-70B-Instruct, and Llama-3.1-8B-Instruct. Our findings reveal a critical trade-off dependent on model scale. While larger models tend to remain socially responsive and more frequently engage with distractor topics, the smaller model shows rigid goal adherence by resisting and rejecting distraction. Although redirection is the most common initial response, subsequent trajectories di-verge substantially. The inclusion of explicit dialogue state demonstrates model-dependent effects, acting as a stabilising anchor for smaller models but providing limited benefit for larger ones. These results suggest that maintaining goal alignment in role-playing agents requires explicitly managing the trade-off between conversational responsiveness and goal adherence.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking

BluePRINT is introduced, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module, and Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory.

Si-Yu Chen, Hao-Ran Wang, Xiaojian Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

How a Chatbot's Response Style Shapes a Classroom: A Multi-Agent Simulation of Students Consulting AI

A virtual classroom of 20 student agents who interact through rule-based chats, quarrels and consultations with friends and, when stressed, may instead consult a counselor AI under one of six style prompts, specifying the agent dynamics completely and discussing the limits of an LLM as generator of state updates.

Rin Tamai, Yu-Ya Dan · 0 citations
2026

Green Bots versus Red Bots: Evaluating Large Language Models for Simulating Persuasion Dynamics in Online Influence Campaigns

A dual-level evaluation framework to assess LLM-based agents at both the individual and collective levels is proposed, finding that while agents capture broad partisan orientations, they underestimate within-group variability and reproduce stereotypical ideological biases.

M. Al Ali, Filip Mihai Muntean, Lucia Donatelli et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Do Personality-Tuned LLMs Make Better Social Agents?

This work investigates whether personality-aware fine-tuning can reduce the consistency and controllability of personality-conditioned dialogue generation compared with instruction prompting alone, and indicates that fine-tuned models are not better at role-playing different personalities than their respective baseline...

Tim Krabbe, Xiao-Dan Shi · 0 citations
#natural language process... Preprint Oct 2026

RELATE: An Evaluation Framework for measuring Relational Orientation of Large Language Models

Large language models (LLMs) are increasingly used for emotional support, raising concern that sustained use may draw users away from their real-world relationships. Yet existing evaluations primarily focus on the safety, empathy, or helpfulness of responses, leaving under-examined a relational question: where does the...

Shivam Shukla, Jihye Kim, Shubham Gaur et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.