Skip to content
Preprint

Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

SuperScout is presented, which routes after scouting the repository: a 7B searcher, SuperScout-7B, first explores the repository and produces a structured handoff whose reproduction claims are sandbox-verified, with false claims stripped before delivery.

Abstract

Frontier language models can resolve repository-level software issues, but each attempt is expensive, and existing routers select a model from the issue text alone. We present SuperScout, which routes after scouting the repository: a 7B searcher, SuperScout-7B, first explores the repository and produces a structured handoff whose reproduction claims are sandbox-verified, with false claims stripped before delivery. The searcher's hidden states, together with the task text, then feed a resume-based router that dispatches the task to one of four frontier fixers. Adding a new fixer requires no retraining. On the full Python slice of SWE-bench Pro (266 tasks) under the benchmark's official capped budget tier, SuperScout matches the best single model's solve rate (159 of 266 for SuperScout, 158 for the best model) at about a fifth of the total cost per solve, and the reported configuration sits above the random traffic-splitting baseline. A no-router ablation, always the cheapest fixer with the handoff, ties the routed system on this benchmark, so the handoff rather than the routing decision carries the result. A paired calibration study points to the mechanism: the handoff appears to redistribute rather than add solving ability, lifting the three cheaper fixers while slightly hurting the strongest, though at $N=99$ the per-fixer effects are directional only; the searcher's hidden states improve cost routing on the calibration labels while the handoff's own text does not. The searcher's compute adds less than half a cent of GPU time per task.

View source

Similar papers

Jul 2026

AllocBench: Measuring Online Tool Allocation Capability in LLM Agents

A paired benchmark that tests whether LLM agents exhibit conscious allocation behavior under a fixed budget in two contexts: an abstract text-based formulation and a code-construction task finds that every frontier model testedacts near-optimally in the abstract framing but fails to transfer this ability to script-writ...

Daniel Wang, Andrew Xu · 1 citation
Preprint Aug 2026

SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks

CRM+RCCR, an architecture-agnostic cost-aware objective that encodes cost preference into continuous relevance targets through per-pair independent scoring, eliminating multi-positive dilution while regularizing queries with similar routing preferences to be closer in the routing space.

Tao Yu, Yi-Fei Qu, Zhi-Qing Cui et al. · 1 citation
Preprint Aug 2026

Loreley: Repository-Scale Program Evolution with Quality-Diversity Search

This work compares configured Loreley QD, sequential champion editing, and independent root proposals in a matched Zstandard experiment and finds that Sequential had the highest observed 48-job mean and median and established a QD advantage.

Mo Chen · 0 citations
Preprint Jul 2026

Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks

A common measurement protocol and hybrid evaluation of four router implementations across RouterBench, BFCL v4, tau2-bench, and WebArena shows that, under these configurations and controls, observed gains track selected-tier composition more closely than demonstrated task-specific targeting.

Kiran Kumār, Santhoshkumar Saminathan · 0 citations
Preprint Sep 2026

ChurnBench: A Drift-Aware Benchmark Demonstrating That Refresh Scheduling, Not Cache Age, Governs Staleness in Agentic AI

In production, agentic systems answer questions over data that lives in several places and keeps changing: licenses are reassigned, users offboarded, prices changed, contracts renewed. Existing retrieval benchmarks freeze the data, so they cannot ask whether an agent's answer is still true, only whether it found the ri...

V. Singh, Preeti Priyam, Gautam Bhowmick · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.