Across both tasks, the proposed LLM-based hierarchical framework consistently outperforms direct task-specific control and optimization, end-to-end reinforcement learning, rule-based and RL-based hierarchical coordination, and prompting-only LLM coordinators, demonstrating the value of heterogeneous-context reasoning, hierarchical execution, and continuation-aware policy learning.
Abstract
Coordinating multiple interacting units in complex engineering systems is challenging when system interactions are difficult to model, operational information is heterogeneous, and low-level actions must satisfy strict constraints. We propose an LLM-based hierarchical framework in which the LLM coordinates interacting units based on heterogeneous operational context, while task-specific controllers or optimizers generate executable and constraint-aware actions. We further introduce Continuation-Aware GRPO to capture the consequences of coordination decisions over subsequent control intervals. Rather than judging a decision only by its immediate outcome, the method also evaluates how the system evolves afterward under the current policy. We validate the framework on multi-ramp traffic control and virtual power plant (VPP) energy management, using simplified system models for training and more realistic simulators for evaluation. Across both tasks, the proposed method consistently outperforms direct task-specific control and optimization, end-to-end reinforcement learning, rule-based and RL-based hierarchical coordination, and prompting-only LLM coordinators, demonstrating the value of heterogeneous-context reasoning, hierarchical execution, and continuation-aware policy learning.
Group Planning-aware Policy Optimization (PlanPO) is proposed, a simple yet effective RL method for learning generalizable planning abilities beyond task-specific high-quality behavior patterns that enables agents to actively learn generalizable and deliberate behaviors spanning interaction planning and textual generation from high-quality rollouts, without degenerating into vanilla length minimization.
D. Liang, Liyuan He, Xuan Feng et al.· 0 citations
Multi-agent large language models solve complex tasks by coordinating several policies in a shared environment. However, existing reinforcement learning methods usually optimize each response or trajectory separately, even when several outputs jointly cause one state transition. Consequently, the update unit differs from the action executed by the system. To address this problem, we propose SRPO (Setwise Relative Policy Optimization), which treats the active set the minimal set of outputs consumed by one transition, as one multi-agent action. Specifically, SRPO combines member log-ratios into one cardinality-normalized set ratio, assigns one relative advantage, and clips the set once. This formulation unifies division of labor and joint co-evolution as actions with different set sizes. Experiments on mathematical reasoning and multi-turn search demonstrate one training interface for fixed, mixed, and dynamically routed workflows across four model scales, with the strongest macro-average results among the reported comparisons. Optimization diagnostics further characterize its stability under different event reductions and set sizes.
Sheng-Tian Yang, Zi-Yun Xiong, Yu Li et al.· 0 citations
Civil aviation fleet planning is a typical long-horizon operations research problem, for which the problem formulation and algorithms could play critical roles in either the computational efficiency or the solution quality. However, conventional meta-heuristic approaches have some inherent drawbacks: during the optimization process, the model is unable to "learn" efficiently from infeasible solutions and gradually concentrate on high feasibility regions, leading to an exponential increase in complexity with increasing planning horizon, and their randomness prevents an optimality guarantee on individual instances. In this paper, we introduce a hierarchical multi-agent proximal policy optimization framework to tackle these problems. To effectively handle ultralong episodes, our approach builds up the optimization architecture by reward shaping and environment designing within a base PPO, which is then boosted through multi-agent coordination and hierarchical decomposition. Large language model (LLM)-based agent workflows are deployed to automatically tune hyperparameters toward the specific domains' performances. Validated on realistic flight data from 2019 shared by our airline partner, our results show that the proposed method can reduce total airlines' operational costs—including direct operating cost and capital cost and achieves a computation speedup in comparison with a conventional optimization baseline.
Li-Jing Liu, James M. Shihua, Qi-Yu Yan et al.· MATEC Web of Conferences· 0 citations
HiDiffTIR is proposed, a Hierarchical Difficulty-aware policy optimization framework for multi-turn TIR that consistently improves multi-turn TIR performance and tool invocation accuracy over strong RL baselines, highlighting the necessity of difficulty-aware credit assignment for effective policy optimization in tool-integrated LLM agents.
Yu-Can Guo, Xiaohan Wang, Miao Su et al.· 0 citations
Effective multi-agent coordination requires aligning incentives while adhering to complex requirements. However, real-world systems often impose situational constraints, context-dependent requirements triggered only under specific conditions, which challenge standard Correlated Equilibria (CE) solutions. We propose Situational-Constrained Density-Based Correlated Equilibria (SC-DBCE), a novel concept in Markov Games that formalizes situational constraints as logic implications. To solve this, we introduce Situational-Constrained Correlated Policy Iteration (SC-CPI), a reinforcement learning algorithm employing a smooth Log-Sum-Exp mechanism for constraint optimization. Evaluations on multi-agent games, smart grids, and warehouse robotics demonstrate that SC-CPI consistently outperforms baselines in both equilibrium quality and constraint adherence. To our knowledge, this is the first method learning CE under situational constraints.
Libo Zhang, Zhi-Rui Zeng, Yang Chen et al.· Proceedings of the Thirty-Fi...· 0 citations
This paper develops a multi-agent reinforcement learning-based (MARL) delegation training that enables agents to make sequential delegation decisions while minimizing the total execution cost and introduces two new frameworks for collaboration and delegation in multi-agent systems.
Ziqing Lu, Avinash Mudireddy, Sarra M. Alqahtani et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.