Skip to content

Multi-Agent System Search via Active Substructure-aware Policy Optimization

Sep 2026 · 0 citations · 67 references
Computer Science

TL;DR

ASPO introduces an Adaptive Query-Selection Mechanism (AQSM) that focuses training on queries at the policy's competence boundary: those it can solve but not yet reliably, and introduces substructure-level rewards that measure output-quality gains within each action's descendant subgraph.

Abstract

LLMs enable multi-agent systems (MAS) to tackle complex tasks, but manually designing agent roles, prompts, and communication structures requires substantial expertise and effort. This motivates learning policies that construct query-specific MAS from execution reward. Existing approaches typically train these policies by repeatedly traversing a fixed set of training queries and assigning rewards at the workflow level. However, this overlooks differences in queries'evolving learning potential and obscures which substructures improve solution quality. In this paper, we propose Active Substructure-aware Policy Optimization (ASPO), a RL framework for query-level MAS search. ASPO introduces an Adaptive Query-Selection Mechanism (AQSM) that focuses training on queries at the policy's competence boundary: those it can solve but not yet reliably. A complementary discovery mechanism widens architectural exploration for hard queries, helping distinguish insufficient exploration from operator capability limits. Beyond query selection, ASPO introduces substructure-level rewards that measure output-quality gains within each action's descendant subgraph. These rewards guide proximal policy optimization to reinforce useful architectural refinements and discourage redundant or harmful computation. Together, these mechanisms prioritize learnable queries and provide fine-grained feedback for learning effective MASs. Across six benchmarks spanning mathematical reasoning, general question answering, and code generation, ASPO ranks first on every benchmark against twelve baselines.

View source

Similar papers

#machine learning Preprint Oct 2026

Dynamic Harness Search: Building Multi-Agent Systems Per-Query via Prediction

Agent harnesses specify the roles, instructions, tools, and communication structure used to solve a task, and the right harness depends on the query. Because the value of each design choice is observable only through execution, tailoring a harness to each query has required either executing alternatives at inference ti...

Som Sagar, Sha-Sha Li, He-Jie Cui et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization

Designing dense reward functions for low-level reinforcement learning (RL) control remains difficult. Recent work uses large language models (LLMs) to iteratively generate and refine reward functions using policy-training feedback within scripted search algorithms. However, evaluating each candidate requires a full RL...

Ming-Hao Li, Rui Tan, Rui-Hang Wang · 0 citations
#natural language process... Preprint Oct 2026

Harness-Search: Guiding Long-Horizon Search through Multi-Agent Coordination

Long-horizon search requires agents to gather evidence across multiple steps and synthesize it into well-supported answers. The recent agent harnesses provide a natural and promising framework to support such long-running search processes. As interaction histories grow, one single agent in harnesses might get stuck and...

Shan-Yong Wang, Zhen-Wen Ji, Lei Jin et al. · 0 citations
#natural language process... Preprint Oct 2026

Programmatic Search Agents: Extending Agentic Search Beyond Query Reformulation

Search agents adapt their queries, yet fixed search interfaces leave candidate processing and evidence presentation outside the agent's direct control. Our trajectory analysis shows that supporting passages can be retrieved yet never delivered to the agent; a same-page oracle intervention shows that changing the return...

Jia-Ming Qian, Hui-Yan Yang, Man-Di Liu et al. · 0 citations
Preprint Aug 2026

PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs

Group Planning-aware Policy Optimization (PlanPO) is proposed, a simple yet effective RL method for learning generalizable planning abilities beyond task-specific high-quality behavior patterns that enables agents to actively learn generalizable and deliberate behaviors spanning interaction planning and textual generat...

D. Liang, Liyuan He, Xuan Feng et al. · 0 citations
#natural language process... Preprint Sep 2026

MAS-OPD: On-Policy Distillation for Multi-agent Systems

MAS-OPD is presented, where Role-Advantage Specialization defines the role advantage as the difference between the teacher signals under target and non-target role conditions, and Privileged Attribution for Coordination attributes an interaction conflict to its source and supplies it to the teacher alone as privileged...

Qi-Yong Zhong, Mao Zheng, Ming-Yang Song et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.