2026· IEEE Transactions on Geoscience and Remote Sensing· Vol 64, pp. 4707115-4707115· 0 citations· 54 references
Abstract
Multiple object tracking (MOT) from unmanned aerial vehicles (UAVs) presents significant challenges due to drastic viewpoint changes and ambiguous top-down appearances. The former often result in large, misleading displacements of objects in the image plane, while the latter makes it difficult to distinguish between objects with similar appearances. To address these, we propose D2Track, a novel graph neural network (GNN)-based tracker that focuses on learning Decoupled and Discriminative representations for robust association. Our main contribution consists of two specialized modules. First, we propose the adaptive feature rectification module (AFRM), which effectively decouples the camera’s ego-motion from the object’s true motion. By estimating the global projective transformation between frames, the AFRM generates a decoupled motion feature that provides a more accurate motion cue for the GNN. Second, to address the difficulty of distinguishing objects with similar appearances, we design a KFD feature using the kernel Fisher discriminant (KFD) algorithm. This method projects the initial reidentification (ReID) features into a more discriminative feature space, significantly improving the model’s ability to differentiate between such similar objects. Extensive experiments on the challenging VisDrone, UAVDT, and SportsMOT benchmark datasets demonstrate that our proposed tracker, D2Track, achieves excellent performance, proving the effectiveness of its decoupled motion and discriminative appearance strategies.
Joint detection-and-embedding (JDE) trackers avoid per-detection crop inference by reading identity features for re-identification (ReID) from the detector. The detection-center readout, however, does not use a track prediction when forming the appearance descriptor. This is restrictive in unmanned aerial vehicle (UAV)...
Multi-Object Tracking (MOT) remains challenging due to object occlusion, complex motions, and detection unreliability in crowded scenarios. We propose an enhanced MOT framework integrating and optimizing state-of-the-art components, specifically Improved Detection Confidence Boost (IDCBoost) and Track-Perspective-Based...
Trung Nghia Huynh, Chi Nhan Huynh, Jia-Ching Wang et al.· International Conference on...· 0 citations
A spatio-temporal regression architecture combining 3D CNNs and attention mechanisms that regresses these fused representations into a relative position estimate and demonstrates the effectiveness of the method on the TII Drone Racing and UZH-FPV datasets.
Nilda G. Xolo-Tlapanco, J. Martínez-Carranza· Unmanned Systems· 0 citations
The results validate the effectiveness of the method in addressing the specific challenges of UAV-based object detection, offering a balanced solution for accuracy and efficiency in resource-constrained scenarios.
Chenguang Zhang, Yangming Guo, Jian-Long Yu et al.· International Conference on...· 0 citations
Multi-object tracking (MOT) has advanced rapidly in urban surveillance and autonomous driving, yet many trackers rely on ReID- and transformer-based appearance encoders and are designed for standard FoV cameras. These assumptions break down for low-cost omnidirectional deployments, where equirectangular projection intr...
Xin Shu, Meegan Gower, Y. Buckley et al.· 0 citations
Small-object detection in unmanned aerial vehicle (UAV) imagery remains challenging due to small target scales, dense target distributions, complex backgrounds, and limited onboard computational resources. To address these coupled problems, this paper proposes ASE-YOLO11, a lightweight improved detector based on YOLO11...
Shi-Hua Hou· 2026 7th International Confe...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.