Skip to content

Author

Xiaomin Li

We have 7 of 25 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Review Oct 2026

AgentPersonaBench: Benchmarking Persona-Driven User Simulation

We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b...

Jin-Tao Huang, Yi-Fan Wang, Hong-Yuan Shen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback...

Fang Wu, Dan-Lei Xing, Yan-Jie Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning, is introduced, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.

Dongwon Jung, H. Ramesh, Yi-Fan Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Certified Long-Horizon Code Agent Evolution via Validation-Gated Skill Optimization

The concept of in-context self-evolution is formalized and VALVE, a validated-gated framework for long-horizon skill optimization is introduced, which establishes finite convergence, provides theoretical guarantees for future-task gain and drawdown, and derive the validation and evaluation holdout sizes required for a...

Yi-Fan Wang, Hao Cheng, Xiao-Min Li et al. · 0 citations
Review Aug 2026

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

MatrAIx is introduced, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users and provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.

Xiaomin Li, Yuexing Hao, Jian Hou et al. · 1 citation
Preprint Aug 2026

CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model Trajectories

Autonomous computer-use agents are increasingly applied to long-horizon tasks requiring coordinated application calls, persistent state tracking, and verifier-sensitive writes, yet they remain prone to procedural failures: misreading application state, tool semantics, or task progress. Procedural memory promises more c...

Zheyuan Deng, Bing-Hang Lu, Han-Qi Feng et al. · 0 citations
Preprint Jul 2026

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do not test whether progress is distributed differently across task difficulty. We find that...

Hanwen Xing, Pengyu Wang, Bingxu Meng et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.