Skip to content

Investigating Assistant Bias in LLM User Simulators Using a Role Vector

Sep 2026 · 0 citations · 77 references
Computer Science

TL;DR

These findings provide a representation-level analysis of LLM user simulators, confirming that assistant bias is structurally identifiable and that user behavior can be directionally analyzed.

Abstract

LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, these simulators exhibit"assistant bias,"a tendency to cooperate and pursue task goals. They rarely reproduce the frustration or disengagement that real users exhibit, compromising evaluation validity. Prior work outlines that this bias is baked in during model training, which role-playing prompts fail to override. We analyze this bias from model activations, extracting a user role vector by contrasting how the model represents user versus assistant perspectives on the same dialogue. We observe two findings: (i) the user direction is identifiable in activations, elicits user-like behaviors, and captures characteristics distinct from assistant traits; and (ii) although user-role activation associates with simulation realism and steering strengthens it, it can exaggerate user behaviors and override individual user profiles. Together, our findings provide a representation-level analysis of LLM user simulators, confirming that assistant bias is structurally identifiable and that user behavior can be directionally analyzed.

View source

Similar papers

#natural language process... Preprint Oct 2026

CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

Recent benchmarks rely on user simulators to evaluate AI agents in multi-turn interaction. While existing simulation techniques demonstrate surface fidelity to human style and behavior, ecologically valid interactive benchmarking also requires alignment in when and how agents fail across simulated and real user populat...

Anjali Kantharuban, Jonas Mueller · 0 citations
#artificial intelligence Preprint Sep 2026

A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation

A three-tier persona vector with 23 operationalized dimensions, demonstrating the model faithfully reproduces real-world difficulty distributions and seven rule-described trait correlations produce auditable co-occurrence patterns without requiring learned covariance matrices.

Rahul Khedar, Eshita, S. R. Thondapu et al. · 0 citations
#artificial intelligence Review Oct 2026

AgentPersonaBench: Benchmarking Persona-Driven User Simulation

We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b...

Jin-Tao Huang, Yi-Fan Wang, Hong-Yuan Shen et al. · 0 citations
#machine learning Preprint Sep 2026

Simulating Disengaged Students to Evaluate LLM-based Tutors

Simulated students generated by computational models provide a practical way to evaluate tutoring strategies and pedagogical approaches used by human and AI tutors. However, such simulations should account for disengaged behaviors, including gaming the system, wheel-spinning, and off-task behavior, because tutors may n...

Xiang-Hui Meng, Jiong-Hao Lin · 0 citations
#natural language process... Preprint Sep 2026

On the Behavioral Traits of LLM Agents

This work proposes A-B-D to infer traits bottom-up from behavioral data (B-data), namely how agents act on their environment and communicate with users, as recorded in existing trajectories, and offers a new lens for understanding AI personality.

Hao-Kai Zhao, Jie Gao, Yunze Xiao et al. · 0 citations

Related blog posts

Microsoft Research Blog Jul 8, 2026

Flint: A visualization language for the AI era

Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, human-editable specifications. The post Flint: A visualization language for the AI era appeared first on Microsoft Research.

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.