Skip to content

Author

Wen-Hao Huang

We have 8 of 25 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems

Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world requests are often open-ended, casually specified, and context-dependent, requiring models not only to follow explicit instructions but also to infer unstated needs from user background and situational context. We introduce xDailyBench, a benchmark of 248 carefully curated tasks spanning 51 scenarios across personal life, white-collar work, learning and research, and cross-domain activities. The tasks are grounded in requests that users have actually completed or genuinely intended to accomplish with AI, and are evaluated with fine-grained binary rubrics covering both explicit and implicit requirements. We evaluate 11 frontier models under standardized agentic settings. The best models achieve a task-level score of 75.6\%, while all models perform substantially worse on implicit than explicit requirements, with gaps no less than 9 percentage points. These results reveal implicit requirement inference as a persistent bottleneck for reliably satisfying real-world everyday user needs.

Yong Peng, Qing-Shui Gu, Li-Ya Zhu et al. · 0 citations
#natural language process... Preprint Sep 2026

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

It is found that generated harnesses remain substantially behind mature human-engineered references on code and on search and research, while matching or exceeding the selected references on writing and machine-learning experimentation, with large variation in execution cost.

Yu-Hao Wu, Jingyuan Zhang, Jia-Jun Shi et al. · 2 citations
Preprint Aug 2026

Repo2Skill-Evo: Repository Skills Go Stale in Silence

Repo2Skill-Evo casts each release transition as a skill-maintenance task: given a V1 skill set and the official V1-to-V2 patch, an agent must update obsolete skill content while preserving guidance that remains valid.

Chenyuan Duan, Ge Shi, Zineng Mao et al. · 0 citations
Preprint Aug 2026

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

This work introduces Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads.

Zining Huang, Haoran Que, Hongxia Zeng et al. · 1 citation

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents

The introduction of MM-BrowseComp, a novel benchmark comprising 400 challenging, hand-crafted questions designed to evaluate multimodal retrieval and reasoning capabilities, is introduced, establishing MM-BrowseComp as a rigorous new standard for the field.

Shilong Li, Xingyuan Bu, Wenjie Wang et al. · 37 citations · ⚡7
#natural language process... Preprint Aug 2026

Aspire: Can Models Self-Evolve from Vague Goals?

This work introduces ASPIRE, a benchmark for vague-goal-driven self-evolution and shows that vague goals redirect search effort toward goal interpretation, and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals.

Yu-Hao Wu, Jingyuan Zhang, Jia-Jun Shi et al. · 0 citations
#natural language process... Preprint Aug 2026

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

These findings show that recognizing successful actions is insufficient; agents must also transform feedback into executable and transferable policies, and provide a unified framework for diagnosing this process and identifying the bottlenecks that prevent agents from translating interaction experience into reliable self-improvement.

Jia-Jun Shi, Siyang Tao, Yu-Hao Wu et al. · 1 citation
Preprint Aug 2026

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

This work systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains and establishes StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.

Li-Ya Zhu, Xin Ma, Tao Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.