Skip to content

Author

Xue-Bing Li

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

Task-Guided Multi-UAV Cooperative Multi-Target Tracking with Gaussian Process-Based Value Correction

The cooperative tracking of multiple ground targets by multiple UAVs remains challenging under partial observability, limited communication, and obstacle constraints, owing to complex target association, difficult task handover, strong coupling among low-level continuous control decisions, and unstable critic value estimation. To address these issues, this paper proposes a hierarchical-guidance and Gaussian-process-corrected multi-agent proximal policy optimization method, termed HGP-MAPPO. Built upon the centralized-training and decentralized-execution paradigm, HGP-MAPPO introduces low-frequency task-guidance signals derived from target-association information, task handover and recovery cues, task priorities, and desired observation geometry. These guidance signals are incorporated as conditional inputs into the low-level actor–critic framework, thereby reducing the policy learning difficulty in jointly handling target tracking, occlusion recovery, obstacle avoidance, and smooth control. Moreover, to alleviate local estimation bias in the neural-network critic under complex partially observable conditions, a Gaussian-process-based residual correction mechanism is designed. Specifically, the posterior mean is used to compensate for value residuals, while the posterior uncertainty adaptively regulates the correction intensity, improving the stability of value evaluation and policy optimization. A sparse inducing-point approximation is adopted to control the training-stage computational cost, while the Gaussian-process module is removed during decentralized execution and, therefore, introduces no additional online inference overhead. Experiments are conducted in standard-obstacle and densely obstructed multi-UAV multi-target tracking scenarios, with DDPG-MHSA, MAPPO, MADDPG, and MATD3 adopted as baselines. The experimental results demonstrate that HGP-MAPPO achieves faster training convergence, higher average episode rewards, and improved target retention rates. It also effectively reduces UAV–target distance fluctuations and the mean absolute temporal-difference (TD) error. Ablation studies further confirm the contributions of task-guidance signals, Gaussian-process residual correction, and uncertainty-aware weighting to cooperative tracking performance and training stability.

Wei Li, Xin Chen, Xue-Bing Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.