We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b...
Jin-Tao Huang, Yi-Fan Wang, Hong-Yuan Shen et al.· 0 citations
Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback...
Fang Wu, Dan-Lei Xing, Yan-Jie Huang et al.· 0 citations
ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning, is introduced, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.
Dongwon Jung, H. Ramesh, Yi-Fan Wang et al.· 0 citations
The concept of in-context self-evolution is formalized and VALVE, a validated-gated framework for long-horizon skill optimization is introduced, which establishes finite convergence, provides theoretical guarantees for future-task gain and drawdown, and derive the validation and evaluation holdout sizes required for a...
Yi-Fan Wang, Hao Cheng, Xiao-Min Li et al.· 0 citations
MatrAIx is introduced, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users and provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
Xiaomin Li, Yuexing Hao, Jian Hou et al.· 1 citation
Autonomous computer-use agents are increasingly applied to long-horizon tasks requiring coordinated application calls, persistent state tracking, and verifier-sensitive writes, yet they remain prone to procedural failures: misreading application state, tool semantics, or task progress. Procedural memory promises more c...
Zheyuan Deng, Bing-Hang Lu, Han-Qi Feng et al.· 0 citations
Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do not test whether progress is distributed differently across task difficulty. We find that...
Hanwen Xing, Pengyu Wang, Bingxu Meng et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.