Skip to content
Review

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

Aug 2026 · 4 citations · ⚡ 1 influential · 43 references
Computer Science

Abstract

We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement is itself a task, and completing one evolution cycle can schedule the next. In experience-driven core evolution, ordinary work and social interaction expose bugs, rough edges, and inefficient context construction that lead to reviewed structural changes. On Terminal-Bench 2.1, an Opus 5 run scores 86.74%, the best result reported on the benchmark. On OSWorld-Verified, an Opus 5 run reaches 90.69%, exceeding the best previously reported score. A five-rollout CL-Bench campaign achieves a normalized reward of 0.2301, setting a new state of the art. Hope is the longest-running publicly documented Ouroboros deployment. It is a 161-day living agent experiment in free evolution under governed human communication across seven surfaces. Human interaction surfaces faults and generates proposals, but the agent decides which changes to pursue. Because a self-developing agent may rewrite its own code and select new model APIs, operational safety becomes a primary design problem: guardrails must remain authoritative under evolutionary and public social pressure. Benchmark campaigns use frozen system snapshots, while Hope continues live evolution on a separate lineage.

View source

Similar papers

Preprint Aug 2026

Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?

This work introduces Open-Ended Optimization (OEO), which keeps the objective, permitted interactions, resource budget, data boundary, and evaluation fixed while allowing the optimizer to compose the improvement process online.

Xue Hui, Fan Yang · 1 citation
#artificial intelligence Preprint Sep 2026

RobustSGPO: Search-Space Control for Agent Harness Evolution

Semantic-gradient-based prompt optimization (SGPO) improves agent harnesses using execution feedback, but its local update rule leaves the choice of edit scope and operation unresolved. We introduce RobustSGPO, which specifies the requested edit, constructs and checks the patch, and continues search from either the inc...

Zi-Bo Zhao, Ji-Jun Shi, Mo-Qing Zhou et al. · 1 citation
#artificial intelligence Review Aug 2026

AutoScientist-Quant: Self-Evolving Coding Agents for Automatic Research in Quantitative Investment

AutoScientist-Quant, a self evolving search process that regards quantitative research as one budgeted search problem, is presented, a self evolving search process that regards quantitative research as one budgeted search problem.

Zong-Qian Li, Yaoyiran Li, Yao-Hui Guo et al. · 2 citations
#natural language process... Preprint Aug 2026

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

This work proposes the Evolutionary Markov Hypergraph Attack (EMHA), a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates, and establishes OpenART as a scalable foundation for studying agent safety in complex, evolving en...

Yunhao Chen, Xin Wang, Yi-Xu Wang et al. · 0 citations
#natural language process... Preprint Sep 2026

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

This work introduces early outcome prediction, a complementary axis of efficiency that instead cuts cost within each task within each task, and instantiates EarlyEval, a lightweight framework that trains a pair of LightGBM success and failure classifiers over behavioral, textual, and reference-solution features and hal...

Yu-Ling Shi, Zhensu Sun, Jun-Sen Dong et al. · 2 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.