Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This...
Ji-Xuan Chen, Jia-Xin Zhang, Qinyuan Ye et al.· 0 citations
Opera is presented, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved, and achieves the highest mean resolve rate among competitive critic baselines on all three benchmarks.
Kai Mei, Zhi-Yuan Hu, Yu-Tong Dai et al.· 0 citations
Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screenshot is only a lossy rendering of the underlying program state, e.g., the files, application backends, and DOM that hold the task data. Different states can produce the s...
Yan Yang, Xiangru Jian, Ziyang Luo et al.· arXiv.org· 0 citations
DarwinX is introduced, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence sh...
Yifang Zhang, Yutong Dai, Juntao Tan et al.· 2 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.