We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b...
Jin-Tao Huang, Yi-Fan Wang, Hong-Yuan Shen et al.· 0 citations
The integration of graphs with Goal-Conditioned Hierarchical Reinforcement Learning (GCHRL) has received increasing attention, as graphs naturally encode task hierarchies for effective subgoal sampling. However, existing methods often overlook intrinsic connectivity information, failing to fully leverage the underlying...
Shuyuan Zhang, Zi-Han Wang, Xiao-Wen Chang et al.· 0 citations
EvoGenUI-Bench is introduced, a benchmark for multi-turn interface maintenance comprising 150 five-turn tasks and 750 turns across three scenarios: information presentation, executable interaction, and tool-grounded external state.
Yue Peng, Lan-Ke Xia, Zi-Han Wang et al.· 0 citations
This work introduces the first benchmark suite and simulation framework tailored to pediatric SIC training, and proposes SIC-Agents, a self-improving framework that generates a clinician-editable skill document to guide simulator behavior.
Zi-Han Wang, A. Slominska, Rennie Bimman et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.