Skip to content

CoEvo: Oracle-Grounded Self-Evolution of a Single Model for Multi-Step Causal Reasoning

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

CoEvo is an oracle-grounded self-evolution framework where a single model alternates between Proposer and Solver, and enables an 8B LLM to sustain self-evolution, surpassing distillation baselines and the strongest proprietary reference on path correctness.

Abstract

Multi-step causal reasoning requires chaining inferences where each step constrains the next. An early error propagates silently, and a correct answer reached via flawed logic evades outcome-level detection. In specialized domains, teacher LLMs err on intermediate steps, safety constraints restrict cloud distillation, and shifting conditions demand adaptation, leaving self-evolution as the practical route. Naive self-evolution can collapse: outcome-only rewards let the model exploit distributional shortcuts, and weak self-evaluation reinforces spurious paths into stable failure patterns. We exploit a key asymmetry: generating a correct chain is hard, but verifying a single step is easy. Many high-stakes domains admit a deterministic, queryable oracle, a physics simulator or rule engine over codified constraints. It checks asserted steps without teacher-level ability and abstains beyond its rules; it can check what the model asserts, never replace it. This enables CoEvo, an oracle-grounded self-evolution framework where a single model alternates between Proposer and Solver. As Solver, the model generates competing chains; intra-group debate exposes disagreement steps, a proxy for the capability boundary, and the oracle adjudicates them into process-level supervision. As Proposer, the same model constructs progressively harder scenarios inside oracle constraints, steering the curriculum toward deep multi-hop chains. Both roles are updated jointly, so training pressure co-evolves with the model. On industrial, clinical, and legal multi-step causal reasoning benchmarks, CoEvo enables an 8B LLM to sustain self-evolution, surpassing distillation baselines and the strongest proprietary reference on path correctness (82.1% vs. 71.4%). The trained model generalizes to unseen categories and systems, preserving root-cause accuracy.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Are Stated Reasoning Steps Causally Load-Bearing?

This work uses synthetic multi-hop lookup tasks to measure faithfulness causally at the activation level, specifically on self-generated reasoning, and aims to measure faithfulness causally at the activation level, specifically on self-generated reasoning.

Abhiram Bhupatiraju, Rayan Nyaupane · 0 citations
Conference Open access Sep 2026

I-EDI: Robust Self-Evolution Agents via Verifiable Counterfactual Simulation

This work argues that robust evolution implies Structural Invariance: a reasoning path is valid only if its core dependency graph remains isomorphic under counterfactual perturbations, and enforces this with Verifiable Counterfactual Simulation (VCS).

Run-Ze Fan, Yong Li · 0 citations
#artificial intelligence Preprint Sep 2026

DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems

Large language model (LLM)-based multi-agent systems have experienced rapid growth in recent years. Despite their promise, such systems remain fragile, frequently exhibiting reasoning and coordination errors that can lead to system-level failures. Failure attribution in such systems relies on tracing natural language i...

Ze-Hao Wang, Lanjun Wang, Shi-Long Jin et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning

Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verifi...

Xing Han, Yu-Xin Wang, Chen Chen et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.