Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This...
Ji-Xuan Chen, Jia-Xin Zhang, Qinyuan Ye et al.· 0 citations
Opera is presented, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved, and achieves the highest mean resolve rate among competitive critic baselines on all three benchmarks.
Kai Mei, Zhi-Yuan Hu, Yu-Tong Dai et al.· 0 citations
We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcement learning using Group Relative Policy Optimization (GRPO). Salesforce Koa is trained on public and synthetically generated data, with no customer data, to improve tool...
Zi-Xiang Chen, Su-Feng Niu, Ying-Chieh Liu et al.· 0 citations
This work formalizes fork placement as locating the pivots of the chain's value curve, where the expected outcome turns, and proposes belief-shift branching, which read the model's answer belief at candidate boundaries and fork just before the step where consecutive beliefs diverge most.
Bin Lei, Yu Li, Prafulla Kumar Choubey et al.· 0 citations
This work introduces RAPTOR - a Role-Aware Private Training framework, which alternates shared and expert optimization and targets each failure directly, using expert-specific clipping and noise together with a public expected-owner denominator and a count-independent update schedule that avoids conditioning on private...
Duc Dm, Khai Le-Duc, D. Nguyen et al.· 0 citations
Random Attention keeps the prompt and evicts uniformly at random within each attention head, computing no score at all; across four models and six reasoning tasks, it matches the strongest baseline in task performance while delivering 32-43% higher throughput than that method when deployed with vLLM.
Heng-Yi Wang, Jie-Lin Qiu, Wenting Zhao et al.· 0 citations
It is suggested that inference-time harness design is an important complement to conventional training-time distillation, enabling strong models to transfer cognitive structure to weaker models without retraining.
Cheng Qian, Wenting Zhao, Liang-Wei Yang et al.· 5 citations
ST-Evidence is introduced, the first human-verified benchmark for both discriminative and generative pixel-level grounding, and scalable, automated generation pipelines are developed to create ST-Evidence-Instruct, a 160k-scale dataset bridging high-level reasoning with fine-grained grounding.
Shijie Wang, Honglu Zhou, Ziyang Wang et al.· arXiv.org· 0 citations
DarwinX is introduced, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence sh...
Yifang Zhang, Yutong Dai, Juntao Tan et al.· 2 citations· ⚡1
The LZ penalty is introduced, a penalty specialized for reducing degenerate repetitions in autoregressive language models without loss of capability and without instances of degenerate repetition, and enables state-of-the-art open-source reasoning models to operate with greedy decoding without loss of capability and wi...
Antonio A. Ginart, Naveen Kodali, Jason Lee et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.