Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This...
Ji-Xuan Chen, Jia-Xin Zhang, Qinyuan Ye et al.· 0 citations
This work formalizes fork placement as locating the pivots of the chain's value curve, where the expected outcome turns, and proposes belief-shift branching, which read the model's answer belief at candidate boundaries and fork just before the step where consecutive beliefs diverge most.
Bin Lei, Yu Li, Prafulla Kumar Choubey et al.· 0 citations
Memory-augmented speculation yields a 19--39\% relative accuracy improvement on action prediction and up to a $2.5\times$ increase on observation prediction tasks with repetitive action spaces as memory accumulates and generalize across speculator models of varying cost.
Yu Li, Qinyuan Ye, Prafulla Kumar Choubey et al.· arXiv.org· 2 citations· ⚡1
A comprehensive re-evaluation of two memory-based methods for self-improving agents is conducted, broadening the scope of evaluation along two axes and hypothesizing that task and environment underspecification contribute to this fragility.
Qinyuan Ye, Yu Li, Yada Pruksachatkun et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.