Skip to content
Preprint

Evidence Before Expansion: Reuse, Spawn, or Defer in Lifelong Expert Pools

Aug 2026 · 0 citations · 6 references
Computer Science

TL;DR

This work proves finite-time anytime validity for the observable surrogate discrepancy of a predictable discriminator sequence, and an unconditional one-sided transfer to the population quantity in which each side's slack is the excess risk of a single discriminator.

Abstract

Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing expert for arriving data, spawn a new one, or defer. We present a decision layer that makes all three outcomes statistically meaningful. Reuse and spawn are posed as one-sided sequential hypotheses on a conditional (mechanism-level) discrepancy, separated by an indifference zone; defer is exactly the state in which neither betting e-process has accumulated sufficient evidence. We prove finite-time anytime validity for the observable surrogate discrepancy of a predictable discriminator sequence, and an unconditional one-sided transfer to the population quantity in which each side's slack is the excess risk of a single discriminator; an empirically observed downward-bias regularity makes the spawn side exactly conservative. Recency without sacrificing the guarantee is obtained by a restarted e-detector: a bank of unwindowed betting supermartingales at geometrically spaced restart times (O(log t) memory), with the error budget spent over restart instances, which preserves lifetime anytime validity; spending over expert-creation order likewise controls multiplicity for unboundedly many experts. On synthetic multi-concept streams, Electricity, Covertype, and the recurrence-heavy INSECTS benchmark, the instance-accounted restarted bank achieves zero false spawns and zero false reuses after switches and matches or exceeds the retired windowed heuristic (INSECTS-reoccurring accuracy 0.675), making the deployed algorithm and the guaranteed algorithm one and the same.

View source

Similar papers

Jul 2026

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

This work casts per-step rollout collection as a budget-constrained sequential allocation problem and introduces SARA (Sequential Adaptive Rollout Allocation), a two-threshold, SPRT-style rule that commits effective groups, abandons saturated ones after a short probe, and reallocates the freed budget to fresh prompts,...

Pixel Nomand, Elena Voss, Marcus Hale et al. · 0 citations
Preprint Aug 2026

Split the Labor: Separating Evidence Interpretation from Decision Aggregation

This work proposes a four-field evidence tuple (hypothesis, reliability bucket, rationale, provenance, provenance) and shows that fixing it determines both halves of the framework and applies to a class of rules beyond language models: score-summing triage engines, diagnostic panels scored by counting positives, and ad...

Zhe-Lun Wu · 1 citation
Jul 2026

Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating

A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment an item arrives, from the past (StreamingLLM, H2O) or from a guess about the future (SnapKV). We recast the choice as an estimation problem on a hidden signal, whether an item will...

Maruthi Vemula, Neeraja Gajula · 0 citations
Preprint Sep 2026

From Experiments to Decisions: Reusing Evidence in Autonomous Coding Research

Autonomous coding agents can remember an experiment yet carry forward a conclusion it does not justify. We reconstruct how evidence is reused in a 400-task NeuroGolf campaign, with selected wellbore-prediction records from the same operator as cross-domain comparisons. A numerical counterexample exposes an overbroad ex...

Bo-Da Cheng · 0 citations
#artificial intelligence Preprint Sep 2026

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

Reasoning traces of large language models are widely read as containing"breakthrough"moments and early-legible fates. Both readings rest on measurements missing a counterfactual control at the level of the claim; we supply both controls. First, a restart-controlled truncation probe separates when a solution fits the co...

Yigit Utku Bulut · 1 citation
Preprint Aug 2026

Loreley: Repository-Scale Program Evolution with Quality-Diversity Search

This work compares configured Loreley QD, sequential champion editing, and independent root proposals in a matched Zstandard experiment and finds that Sequential had the highest observed 48-job mean and median and established a QD advantage.

Mo Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.