Time series agents answer analytical questions by calling external tools, and which tools they carry is decided by people before the agent runs. However, we identify two failures in this setup. Human-Agent Tool Misalignment: a library of 21 expert-curated tools helps on some tasks and hurts on others, dropping anomaly...
Jie Yang, Yan Zheng, Jia-Rui Sun et al.· 0 citations
The evaluation of large language models (LLMs) on coding tasks has primarily focused on performance metrics such as pass@k. As LLMs continue to advance, many models now meet baseline performance requirements, reducing the discriminative power of performance-based evaluation alone. Yet a key question remains largely une...
Jun-Peng Wang, Yuzhong Chen, Menghai Pan et al.· 0 citations
Bazaar is introduced, a dynamic sealed-bid benchmark for multi-attribute auction under multi-attribute auction under these conditions, grounded in closed-form customer utilities, enabling exact evaluation.
Shimaa Ahmed, Yiwei Cai, Mohsen Minaei et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.