Claw-SWE-Bench is introduced, a unified benchmark that enables researchers to systematically assess the capabilities and efficiency of general-purpose harnesses on software engineering tasks and guide the development of more capable and efficient harnesses.
Meng-Yu Zheng, Kai Han, Bo-Xun Li et al.· arXiv.org· 5 citations
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-native models developed to explore this path through agentic post-training. Our system combines a heter...
NeoHorse Team, Guo-Liang Cao, Guo-Hao Dai et al.· 0 citations
The report studies singleton and multi-model routing on agentic benchmarks including DRACO and PinchBench, and argues that agentic routing is not merely cost control, but a data engine for agent-native training.
Xinchen Liu, Hang Zhou, Ying-Jie Zong et al.· arXiv.org· 1 citation· ⚡1
These results support depthwise convolution as a lightweight complement to self-attention for modeling short-range token interactions and suggest that the convolution makes repeated token IDs more sensitive to their immediate context.
Yu-Chuan Tian, Yingte Shu, Wei He et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.