Recursive Synthetic Terminal Tasks (RST) is presented, a recursive verified synthesis framework for constructing long-horizon terminal-agent tasks at scale and shows no ceiling, indicating that the process can continue well beyond the scale reported here.
Zhongzhi Li, Yucheng Shi, Zongxia Li et al.· 5 citations
SkillEval is used to evaluate skills in controlled quality tests and it is used for diagnosing weaknesses in skill documents and guiding targeted revisions, and it is shown that SkillEval reliably distinguishes skills of different quality.
Jia-Hui Han, Qinuo Li, Ziheng Peng et al.· 0 citations
This work introduces Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing, and analyzes failure modes and error patterns to support future progress on long-horizon terminal agents.
Zongxia Li, Zhongzhi Li, Yucheng Shi et al.· arXiv.org· 11 citations· ⚡2
SymBOL accurately recovers governing equations and provides interpretable pathways for equation discovery when applied to real-world systems in materials science and epidemiology, and underscores the potential of SymBOL for advancing scientific discovery.
Jiaxu Cui, Qifei Li, Wei-Ting Liu et al.· IEEE Transactions on Pattern...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.