2026· IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing· Vol 19, pp. 29136-29152· 0 citations· 44 references
TL;DR
DenseRS-CLIP, a region-structured two-stage adaptation framework for enhancing dense representations of remote sensing CLIP models, achieves consistent improvements over representative remote sensing VLFMs and controlled distillation variants, and validate the effectiveness of weakly localized, branch-specific distillation for dense remote sensing vision–language representation learning.
Abstract
Existing remote sensing vision–language foundation models mainly follow the CLIP-style global image-text alignment paradigm. While effective for image-level semantic understanding, this paradigm leaves patch-level representations insufficiently discriminative and spatially inconsistent for dense prediction tasks. To address this issue, we propose DenseRS-CLIP, a region-structured two-stage adaptation framework for enhancing dense representations of remote sensing CLIP models. In the first stage, we perform dual-granularity image-text contrastive pretraining to obtain a domain-adapted CLIP model with robust global semantic alignment. In the second stage, we introduce an attention-decoupled dual-branch distillation framework that reuses existing bounding-box and mask-derived localization annotations to construct ROI-level distillation signals without requiring additional region-text descriptions. Specifically, dense features are decoupled into content and context branches. The content branch is optimized by region-structured semantic distillation with a region correlation constraint, which improves local semantic discriminability and suppresses interregion feature homogenization. The context branch is optimized by DINOv3-guided topology distillation, which aligns patch-level self-similarity structures to improve spatial consistency and boundary awareness. Experiments on seven benchmarks covering region classification, visual grounding, and referring image segmentation show that DenseRS-CLIP achieves consistent improvements over representative remote sensing VLFMs and controlled distillation variants. These results validate the effectiveness of weakly localized, branch-specific distillation for dense remote sensing vision–language representation learning.
A three-stage framework that integrates complementary priors from Contrastive Language–Image Pre-training, Self-Distillation with No Labels version 2 (DINOv2), and the Segment Anything Model (SAM) is proposed, demonstrating the effectiveness and generalizability of the proposed framework across diverse remote sensing s...
Open-vocabulary semantic segmentation (OVSS) of remote sensing faces severe performance degradation when encountering unseen scene distributions caused by geographic, sensor, and resolution variations. Existing vision–language approaches provide strong semantic priors but lack scene-invariant structural representations...
Wu-Biao Huang, Hu-Chen Li, Shuai Zhang et al.· IEEE Transactions on Geoscie...· 0 citations
DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...
Mustafa Alawadi, M. Fateh· Jordanian Journal of Compute...· 0 citations
A context-aware referring expression segmentation model for remote sensing that provides a solution for accurate multi-scale target segmentation guided by natural language descriptions in complex remote sensing scenes.
Open-vocabulary semantic segmentation (OVS) of remote sensing imagery is a challenging pixel-level task requiring strong generalization and adaptation to the spatial characteristics of remote sensing data. Although existing vision-language foundation models perform well in general domains, their image-level classificat...
Chang-Hao Zhao, Ling-Lin Zeng, Hai Liu· 0 citations
The emergence of large-scale vision–language models (VLMs) has significantly advanced remote sensing image–text retrieval (RSITR) by providing powerful cross-modal semantic priors. However, when adapted to the remote sensing (RS) domain, these models struggle to capture fine-grained representations due to their inheren...
Wen-Liang Du, Xiao-Yu Xu, Jia-Qi Zhao et al.· IEEE Transactions on Geoscie...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.