MDF-Det: Motion-Aware Decoupling and Scene Filtering for Wide Area Small Moving Target Detection
Abstract
Highlights What are the main findings? MDF-Det establishes a unified coarse-to-fine framework that jointly addresses weak-target preservation, dense-target decoupling, and scene-induced false-alarm suppression in wide-area motion imagery. The spatial attention-guided decoupling and learned scene-prior filtering mechanisms separate merged target responses and suppress contextually implausible detections, enabling MDF-Det to achieve an average F1 score of 0.878 across six WPAFB AOIs. What are the implications of the main findings? The results demonstrate that reliable WAMI detection requires coordinated control of candidate recall, dense-target separation, and scene-level contextual filtering rather than relying solely on local motion cues. The proposed modular framework provides an interpretable and extensible solution for detecting tiny and densely distributed moving targets under complex wide-area aerial imaging conditions. Abstract Object detection in Wide Area Motion Imagery (WAMI) is crucial for large-scale intelligent surveillance and monitoring systems. However, detecting extremely small moving targets in low-frame-rate grayscale WAMI remains highly challenging. In particular, three key problems limit the performance of existing methods: weak motion responses caused by low target contrast can lead to missed detections; densely distributed targets often produce merged responses that are difficult to separate; and registration artifacts, parallax, and dynamic background clutter can generate numerous false alarms. To address these problems, we propose MDF-Det, a coarse-to-fine spatiotemporal framework for WAMI small moving target detection. The framework consists of three complementary components. First, a Motion and Appearance Feature Fusion (MAFF) strategy integrates dense optical-flow-derived motion saliency with grayscale frame differencing to enhance weak target responses and improve candidate preservation. Second, a Spatial Attention-Guided Target Decoupling (SA-TD) module employs fine-grained heatmap decoding and step-threshold degradation to separate merged responses in densely populated scenes. Finally, a Scene-Prior Guided Filtering (SPGF) mechanism learns complementary vehicle-accessibility and motion-activity priors from scene context to suppress contextually implausible false alarms caused by complex background interference. Extensive experiments on six evaluation AOIs of the WPAFB 2009 dataset demonstrate that MDF-Det achieves an average F1 score of 0.878, corresponding to a relative improvement of 5.5% over the strongest baseline.