Sci-MMR is introduced, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions, and it is found that current answer-centric benchmarks substantially overestimate the evidence-gro...
Jia-Qiang Li, Ya-Jie Yang, Zhi-Heng Xi et al.· 0 citations
A novel textual representation of fault trees is proposed, and a benchmark for multi-turn dialogue systems that emphasizes robust interaction in complex environments is constructed, evaluating a model's ability to assist in malfunction localization.
Yuhui Wang, Zhi-Xiong Yang, Ming Zhang et al.· arXiv.org· 0 citations
ContextWeave is introduced, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams and motivates memory systems that optimize not only retrieval relevance but also reliable use during execution.
Bo Wang, Yu Yao, Enxi Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.