Skip to content
Open access

LDF-Net: a transformer-based lightweight detail fusion network for UAV engineering vehicle detection

Sep 2026 · Journal of King Saud University: Computer and Information Sciences · Vol 38 · 0 citations · 54 references

TL;DR

The proposed dataset and detection framework provide an effective solution for engineering vehicle perception in UAV-based intelligent inspection applications and improve multi-range feature modeling and detail-preserving cross-scale fusion.

Abstract

UAV engineering vehicle detection is a fundamental task for intelligent construction monitoring, as engineering vehicles directly reflect construction progress, equipment deployment, and potential safety risks in dynamic construction scenes. However, reliable detection in UAV construction scenes remains difficult due to three major challenges: the lack of a dedicated benchmark dataset, severe scale variation and background interference, and the loss of discriminative details during cross-scale feature fusion. To address these issues, this paper first constructs a dataset named UAV Engineering Vehicles in Challenging Construction Scenes (UEV-CCS), which covers nine representative categories related to engineering vehicles. Then, a Transformer-based detector named LDF-Net is proposed for UAV engineering vehicle detection. Specifically, a Lightweight Grouped Hybrid Attention (LGHA) block is designed to enhance multi-range feature representation in a lightweight manner, while a Detail-Preserving Contextual Fusion (DpcFusion) block is introduced to preserve informative high-resolution details during cross-scale interaction. Together, these two modules improve multi-range feature modeling and detail-preserving cross-scale fusion, enabling LDF-Net to produce more discriminative representations for small and scale-varied engineering vehicles in UAV scenes. Experimental results on the UEV-CCS dataset show that LDF-Net achieves 92.65% mAP@0.5 and 89.23% Recall. Compared with the best competing model, LDF-Net improves mAP@0.5 by 0.62 percentage points while reducing the parameter count by 68.2%. Comprehensive experiments, including ablation studies and external generalization evaluation on VisDrone2019, further verify the effectiveness and robustness of the proposed method. The proposed dataset and detection framework provide an effective solution for engineering vehicle perception in UAV-based intelligent inspection applications.

Read PDF

Similar papers

Open access 2026

LiteUAV-Det: An Efficient Lightweight Deep Learning Framework for UAV-to-UAV Small Target Detection in Complex Aerial Scenes

Detecting small UAV targets in air-to-air scenarios is important for aerial security, autonomous surveillance, and multi-UAV defense systems. However, this task remains challenging due to severe scale variation, motion blur, and complex aerial backgrounds. Existing detectors suffer from performance degradation on extre...

A. Khan, Muhammad Ali Farooq, Xiao-Feng Bai et al. · 0 citations
Sep 2026

FUS-DETR: a robust algorithm for drone detection from a cluttered background in airborne images

An object detection model, FUS-DETR, specifically designed for UAV target detection in air-to-air scenarios, using a transformer-based architecture and incorporating three key innovations, demonstrating the strong potential of this model for air-to-air UAV detection tasks.

Si-Yuan Duan, Geng Zhang, Xin Li et al. · 0 citations
Open access 2026

WTM-YOLOv11: A Small Target Detection Model for Complex Scenes in UAV Applications

Drones serve as critical tools for urban security surveillance and emergency response. However, aerial object detection faces inherent challenges, as key targets such as pedestrians and vehicles often appear tiny, densely distributed, and heavily occluded. These issues reduce detection accuracy and hinder real-world se...

Xin-Yang Wang, Hong-Wei Zhu, Long-Bo Xu et al. · 0 citations
Open access Aug 2026

HSAR-DETR: Hierarchical Spatial–Frequency Attention Network for UAV Small Object Detection

HSAR-DETR is proposed, a detection framework that jointly improves hierarchical feature representation, cross-scale refinement, and geometry-aware localization and experimental results on the VisDrone, RSOD, and TinyPerson datasets demonstrate improved detection performance.

Cheng Zhang, Zhibo Guo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.