Skip to content
Conference

Cooperative Multi-UAV Target Exploration with Graph-Based Reinforcement Learning

Jul 2026 · Fall Joint Computer Conference · pp. 225-232 · 0 citations · 30 references

Abstract

Unmanned aerial vehicles (UAVs) offer several advantages, including high mobility, flexible deployment, low cost, and strong adaptability to complex environments, making them highly promising for applications such as disaster search and rescue, environmental monitoring, inspection, and reconnaissance. For target exploration tasks in unknown environments, multiUAV systems can expand the search area, improve exploration efficiency, and enhance the robustness of task execution through cooperation, which makes this problem of significant research interest. However, such tasks still face several challenges, including partial observability of environmental information, complex cooperative decision-making, and difficulties in credit assignment among multiple UAVs. Reinforcement learning is capable of learning decision-making policies autonomously through interaction with the environment, providing a new perspective for solving cooperative exploration problems in complex environments. To address these issues, we propose a cooperative decision-making method for multi-UAV target exploration. By incorporating target-related information, the proposed method enhances the cooperative exploration capability of UAVs in unknown environments, while a tailored reward design is adopted to improve the coordination efficiency of multiple UAVs. Experimental results show that the proposed method exhibits strong adaptability to different team sizes and sensor configurations, learns effective cooperative behaviors, and outperforms classical exploration methods across multiple performance metrics, thereby demonstrating its effectiveness in multi-UAV target exploration tasks.

View source

Similar papers

Open access Aug 2026

Efficient Exploration-Enabled Multi-Agent Reinforcement Learning for Multi-UAV Cooperative Target Search

A novel method named AEQMIX is proposed, which integrates trajectory entropy maximization into QMIX, an advanced Multi-Agent Reinforcement Learning (MARL) method, to encourage efficient exploration in multi-UAV Cooperative Target Search.

Peng Chen, Tian-Xu Li, Wei-Xing Xia et al. · 0 citations
Conference Jul 2026

Spatially Predictive Intent for Multi-Agent Coordination in UAV Exploration

Distributed multi-UAV systems play an important role in applications such as search and rescue, disaster response, environmental monitoring, and autonomous reconnaissance. These tasks often require multiple UAVs to coordinate navigation and sensing so as to improve efficiency and expand useful environment coverage. How...

Cheng-Lin Tang, Lei Liu, Xudong Lu et al. · 0 citations
#reinforcement learning Open access Sep 2026

A hierarchical navigation decision-making method for UAV swarms in unknown communication-constrained environments

In recent years, Unmanned Aerial Vehicles (UAVs) have gradually been widely used in various fields such as regional search and disaster relief, and the development of related technologies has also experienced unprecedented growth. Compared to individual UAVs, the collaborative execution of tasks by UAV swarms has more...

Hua-Jie Xiong, Bao-Guo Yu, Jing-Kui Zhang et al. · 0 citations
Open access Aug 2026

Multi-Objective Distributed Task Allocation for UAV Swarms with Limited Interactions

This paper analyzes the factors affecting communication interactions between UAVs and proposes a bidding-based grouping method to eliminate ineffective communication interactions, and introduces a network simplification algorithm based on reducing the number of triangular network topologies to optimize the communicatio...

Wei-Xing Xia, Peng Chen, Fei-Fei Song et al. · 0 citations
Open access Jul 2026

An Experience-Guided MAPPO Framework for Multi-UAV Cooperative Tracking in Continuous Action Spaces

A cooperative guidance law based on the experience-guided multi-agent proximal policy optimization (E-MAPPO) algorithm is proposed for multiple unmanned aerial vehicles (UAVs) to track dynamic points of interest in civilian applications and results indicate that the proposed method generalizes well to different types o...

Hao Xiong, Minghu Tan, Xiaoyu Liu et al. · 0 citations
Open access Aug 2026

A Multi-UAV Cooperative Path-Planning Method for Complex Obstacle Environments

Experimental results demonstrate that the Improved Experience Replay Multi-Agent Deep Deterministic Policy Gradient algorithm outperforms other comparison algorithms in terms of convergence speed, training stability, and path-planning performance.

Long Wen, Hui Tan, Yuxi Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.