We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b...
Jin-Tao Huang, Yi-Fan Wang, Hong-Yuan Shen et al.· 0 citations
The OmniEdit-Bench provides a comprehensive and reliable testbed for evaluating instruction-based video editing and offers insights into future research directions, including an accuracy-aware penalty mechanism that conditions other scores on accuracy, preventing visually plausible but incorrect edits from receiving in...
Chenxuan Miao, Yutong Feng, Yi Lu et al.· 1 citation
This work presents S$^2$GE as an instance showing that diagnosis-driven interface design can improve native decoder usability and introduces an intervention triangle with three matched conditions: readable graph evidence, shuffled graph evidence, and no-graph input that separates evidence inclusion, structural readabil...
Xiao-Yu Guo, Peng-Cheng Chen, Jiong Yu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.