Skip to content
Review

Continuous Improvement and Parallel Autonomous Exploration: An LLM-Agent Framework for Searching Large Solution Spaces

Aug 2026 · 0 citations · 26 references
Computer Science

TL;DR

A framework that gives LLM agents two mechanisms for searching large solution spaces autonomously: a leaderboard scored on held-out data acts as a reward signal that drives each agent to refine its solutions over repeated submissions, and a loop that operates even with a single agent.

Abstract

We present a framework that gives LLM agents two mechanisms for searching large solution spaces autonomously. First, a leaderboard scored on held-out data acts as a reward signal that drives each agent to refine its solutions over repeated submissions, a loop that operates even with a single agent. Second, the framework enables running many agents in parallel, fully autonomously, with no human in the loop: agents independently analyze, survey methods, implement, self-evaluate, submit, and revise, while a moderator agent handles only logistics. Running agents in parallel under the shared reward broadens the explored region of the solution space rather than refining the single seeded paradigm. We instantiate the framework on product-to-catalog matching (a core e-commerce retrieval task with a large, category-structured solution space), posed as selective prediction with a precision-coverage operating point. A single agent refines within its seeded paradigm, whereas parallel autonomous agents surface qualitatively different solutions. On this testbed, best qualified coverage (>=95% P@1 per category) reaches 47.8-57.4% with a single agent and 62.8-69.4% with five, against a 33.3% baseline. Our contribution is the framework itself: a continuous-improvement reward loop and a substrate for fully autonomous parallel exploration, backed by case-study evidence.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Multi-Agent System Search via Active Substructure-aware Policy Optimization

ASPO introduces an Adaptive Query-Selection Mechanism (AQSM) that focuses training on queries at the policy's competence boundary: those it can solve but not yet reliably, and introduces substructure-level rewards that measure output-quality gains within each action's descendant subgraph.

Bei-Cheng Xu, Bo-Wen Fan, Wei Qian et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments

Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them without coordination. We a...

Shengbin Yue, Hongru Wang, Siyuan Wang et al. · 0 citations
#artificial intelligence Review Sep 2026

RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leav...

Zheng-Yu Chen, Lin-Feng Liu, Hong Li et al. · 0 citations
Preprint Sep 2026

Certifying cooperation: a novel approach to cooperative multi-agent task generation

This framework exposes the gap between rewarded partial completion and realized cooperation by certifying what cooperation successful completion requires and using temporal cooperation graphs to reveal what policies exhibit.

Yannick Molinghen, Hugo Charels, Tom Lenaerts · 0 citations
#artificial intelligence Preprint Sep 2026

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orc...

Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel Abbé · 0 citations
#artificial intelligence Preprint Sep 2026

Controlling Collectives of AI Agents in Reasoning Space with Spatial Transformers

Large Language Models (LLMs) introduce an exciting new paradigm for planning and navigation in robotics, but fail on even simple multi-robot tasks as team sizes grow. We propose COMPASS, a scalable, decentralized multi-robot architecture for controlling large collectives of agentic robots with reasoning space feedback...

Frederic Vatnsdal, Roshan Gopal, Romina García Camargo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.