Skip to content

Author

Shihan Dou

We have 7 of 117 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

EvoIn: Bridging Evolution and Internalization for Agent Fine-Tuning

EvoIn is an agent fine-tuning framework that bridges evolution and internalization, and consistently enables agents to learn stronger decision-making procedures, raising the pass rate by 10.9 points in-domain and by 9.2 points out-of-domain.

Shi-Han Dou, Shao-Hua Liu, Zhong-Hang Lu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ExplorationBench: Measuring AI Systems'Exploration in Verifiable Alien Worlds

ExplorationBench turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds, and finds that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can s...

Ming Zhang, Zhen Xiang, Pei-Zhong Gao et al. · 0 citations
Preprint Aug 2026

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address thi...

Junjie Ye, Zhuohui Sheng, Shao-Hua Liu et al. · 0 citations
Conference Open access Sep 2026

MathCritique: Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

A critique-in-the-loop self-improvement method that incorporates critique-based supervision into the actor’s self-training process and improves the actor’s exploration efficiency and solution diversity, especially on challenging queries, leading to a stronger actor model.

Zhi-Heng Xi, Dingwen Yang, Jixuan Huang et al. · 0 citations

JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees

A novel textual representation of fault trees is proposed, and a benchmark for multi-turn dialogue systems that emphasizes robust interaction in complex environments is constructed, evaluating a model's ability to assist in malfunction localization.

Yuhui Wang, Zhi-Xiong Yang, Ming Zhang et al. · 0 citations
#natural language process... Preprint Aug 2026

Agents in the Large: Perception-Centered Architecture for Persistent Agents

Pera describes a persistent agent organized around perception and control components that continually perceive service-relevant signals from episodic task executions, internal context, and changes in the surrounding environment, and use these signals to construct lifecycle tasks.

Shi-Han Dou, Haoxiang Jia, Shichun Liu et al. · 1 citation
Preprint Aug 2026

MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents

Results show that function-aware memory arbitration enables accessible information to guide actions more effectively, and improves post-failure recovery and reduces failed-action repetition and state-action recurrence.

Jiajun Dong, Yutao Hu, Fengrui Fan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.