Small object detection in remote sensing (RS) imagery remains fundamentally challenging due to severe information degradation caused by limited spatial resolution and complex background interference. In deep neural networks, such degradation is further exacerbated by irreversible information loss during conventional downsampling, leading to weak and ambiguous feature representations. To address this issue, we propose a frequency–spatial joint decoupling and adaptive perceptual aggregation module (FSD-APAM), which explicitly separates signal-level detail preservation and perceptual-level feature discrimination within a unified framework. Specifically, a frequency-domain detail decoupling (FDD) unit leverages discrete wavelet transform to construct reversible feature pathways, enabling the recovery of high-frequency edge information suppressed during downsampling. Complementarily, a spatial salience decoupling (SSD) unit introduces a lateral inhibition mechanism to enhance isolated target responses while suppressing structured background interference. To further ensure global contextual consistency with low computational overhead, an adaptive contextual perceptual aggregation (ACPA) unit is designed to facilitate efficient interaction between sparse target cues and dense semantic representations. Extensive experiments on AI-TODV2, LEVIR-Ship, and VisDrone demonstrate that the proposed method consistently improves detection performance across multiple mainstream architectures without requiring substantial architectural modifications. In particular, FSD-APAM achieves superior accuracy in detecting extremely small objects while maintaining competitive efficiency, highlighting its practical value for large-scale RS applications. The source code is available at https://github.com/cskkx1/FSD-APAM
Ying Gao, Zongshuai Zhang, Zheng-Yu Zhu et al.· IEEE Transactions on Geoscie...· 0 citations
Fine-grained visual classification (FGVC) plays a crucial role in the realm of computer vision. Recently, multimodal FGVC methods, leveraging textual descriptions as semantic guidance, have gained considerable attention. However, current approaches often encounter two primary limitations: 1) Redundant or ambiguous textual descriptions: existing methods rely on raw or generated descriptions without filtering, introducing redundant and ambiguous semantic noise; and 2) Underutilization of hierarchical visual features: most approaches align single-layer visual features with auxiliary semantic embeddings, underutilizing hierarchical information. To address these challenges, we propose a task-oriented multimodal FGVC framework that eliminates textual redundancy while enhancing multi-layer alignment between cross-modalities. Specifically, our method comprises two key components: Hierarchical Semantic Purification (HSP) and Multi-layer Cross-Modal Alignment (MCA). The former employs a semantic distillation dictionary to eliminate redundant elements and uses a self-attention mechanism for ranking and semantic refinement. The latter establishes effective cross-modal fusion by integrating multi-layer features with purified text features, effectively combining multi-scale visual representations. Experimental results on 7 public datasets demonstrate that our proposed method outperforms existing counterparts, contributing to advancements in fine-grained visual classification.
Meng-Huan Zhang, Qing Cai, Fan Zhang et al.· IEEE Transactions on Image P...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.