The last mile toward enterprise AGI is a company that runs itself. Training and adapting such agents require longitudinal enterprise data, which remain scarce, costly to acquire, and often restricted by privacy constraints. Historical archives are also frequently incomplete and record only what actually happened. They...
Jing-Ying Zeng, Zhen-Wei Dai, Jin-Ning Li et al.· 0 citations
Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed reliably, yet deployed tools can fail transiently, persistently, or silently. Robust recovery therefore requires more than repeated retries: an agent may need to retry the same path, switch to an alternative, or recognize...
Chao-Ran Chen, Vylinh J. Nguyen, Zi-Ji Zhang et al.· 2 citations
Shopping assistants are shifting from ranked product lists toward structured decision support, where systems must synthesize shopper context, product evidence, and next-step guidance into a coherent recommendation experience. This changes the unit of evaluation: a fluent response can still fail by ignoring shopper cont...
Weimin Lyu, Chen Luo, Guangrui Li et al.· 0 citations
Span-Level Uncertainty Estimation (SLUE) is formalized, a new task that targets the natural granularity for uncertainty: semantically coherent text spans, each conveying a single assessable unit of meaning.
Yimeng Zhang, Yingying Zhuang, Ziyi Wang et al.· arXiv.org· 0 citations
This formulation enables a systematic study of key self-improvement factors through the proposed Evo-Harness, and provides a principled understanding of how LLM agents can effectively learn on the fly.
Tian-Xin Wei, Zhan Shi, Min-hua Lin et al.· 14 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.