Skip to content

GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks

Jul 2026 · arXiv.org · Vol abs/2607.05369 · 2 citations
Computer Science

TL;DR

Graph-as-Policy (GaP) is introduced, a multi-agent coding harness that generates directed computation graphs with perception, planning, and control nodes from a Modular Open Robot Skill Library (MORSL), and can achieve success rates that significantly outperform baselines.

Abstract

For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable robot programming with the open-world adaptability of model-free policies? We focus on"Variational Automation"(VA), a class of tasks that have larger variations in object geometry and pose than fixed automation. Model-free policies often struggle to close the reliability gap for VA tasks, which must be executed persistently and reliably in commercial and industrial applications. Motivated by prior work on Task and Motion Planning (TAMP) and the Robot Operating System (ROS), we introduce Graph-as-Policy (GaP), a multi-agent coding harness that generates directed computation graphs with perception, planning, and control nodes from a Modular Open Robot Skill Library (MORSL). GaP then generates an internal simulation environment to rehearse task instances with different graphs in parallel to iteratively refine the graph structure and parameters to improve success rates and throughput. Evaluation with 8 new open VA task benchmarks, 4 in-simulation and 4 in real-world, suggests that GaP can achieve success rates that significantly outperform baselines. Details, code, and data can be found online: https://graph-robots.github.io/gap

View source

Similar papers

2025

Learning and Planning Multi-Agent Tasks via an MoE-based World Model

M3W is a novel approach that applies mixture-of-experts (MoE) to world model instead of policy, enabling both learning and planning, and demonstrates superior performance, sample efficiency, and multi-task adaptability.

Zi-Jie Zhao, Zhong Zhao, Kaixuan Xu et al. · 9 citations

Teach and Grow: An Agent-Centered Architecture for General Robot Learning

Vision-language-action (VLA) and world-action models typically absorb unfamiliar manipulation tasks through additional robot data collection and policy optimization. This recurring retraining burden slows the acquisition of new behavior. We present Teach-and-Grow Learning (TGL), a training-free architecture that turns...

Chang Nie, Zhe Liu, He-Sheng Wang · 1 citation
Preprint Aug 2026

Revisiting the"Push-T"Robot Manipulation Task with Agentic Robotics

This short paper revisits the Push-T task in the context of emerging advances in Agentic Robotics where an LLM coding agent -- Claude Code with Fable 5 -- is prompted to create an algorithmic solution that does not require any demonstration data.

Shuang-Yu Xie, Kai-Peng Chen, Ken Goldberg · 0 citations
Conference Open access 2025

Multi-Robot Cooperative Path Planning: Theories, Algorithms, and Applications

This paper provides a thorough survey and integrative presentation of cooperative path planning for multi-robot systems operating in dynamic, cluttered, and partially observable environments and proposes research directions including learning-augmented heuristics, unified safety-aware planning, adaptive MPC – CBF filte...

Yun Pan · 0 citations
#artificial intelligence Open access Aug 2026

Generalizable Multi-Agent Planning From Signal Temporal Logic Specifications via Diffusion

A new diffusion method for multi-agent planning with STL specifications is introduced, making the approach generalizable to novel formulas whose predicates are placed anywhere within the goal region covered during training, while achieving the same scalability as existing learning-based methods.

Joe Eappen, Zikang Xiong, S. Iyengar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.