Skip to content
Open access

DenseRS-CLIP: Enhancing Dense Feature Representation of Remote Sensing CLIP via Attention-Decoupled Dual-Branch Distillation

2026 · IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing · Vol 19, pp. 29136-29152 · 0 citations · 44 references

TL;DR

DenseRS-CLIP, a region-structured two-stage adaptation framework for enhancing dense representations of remote sensing CLIP models, achieves consistent improvements over representative remote sensing VLFMs and controlled distillation variants, and validate the effectiveness of weakly localized, branch-specific distillation for dense remote sensing vision–language representation learning.

Abstract

Existing remote sensing vision–language foundation models mainly follow the CLIP-style global image-text alignment paradigm. While effective for image-level semantic understanding, this paradigm leaves patch-level representations insufficiently discriminative and spatially inconsistent for dense prediction tasks. To address this issue, we propose DenseRS-CLIP, a region-structured two-stage adaptation framework for enhancing dense representations of remote sensing CLIP models. In the first stage, we perform dual-granularity image-text contrastive pretraining to obtain a domain-adapted CLIP model with robust global semantic alignment. In the second stage, we introduce an attention-decoupled dual-branch distillation framework that reuses existing bounding-box and mask-derived localization annotations to construct ROI-level distillation signals without requiring additional region-text descriptions. Specifically, dense features are decoupled into content and context branches. The content branch is optimized by region-structured semantic distillation with a region correlation constraint, which improves local semantic discriminability and suppresses interregion feature homogenization. The context branch is optimized by DINOv3-guided topology distillation, which aligns patch-level self-similarity structures to improve spatial consistency and boundary awareness. Experiments on seven benchmarks covering region classification, visual grounding, and referring image segmentation show that DenseRS-CLIP achieves consistent improvements over representative remote sensing VLFMs and controlled distillation variants. These results validate the effectiveness of weakly localized, branch-specific distillation for dense remote sensing vision–language representation learning.

Read PDF

Similar papers

Open access Aug 2026

Weakly Supervised Remote Sensing Segmentation via Decoupled Cross-Modal Distillation and Semantic-Guided Refinement

A three-stage framework that integrates complementary priors from Contrastive Language–Image Pre-training, Self-Distillation with No Labels version 2 (DINOv2), and the Segment Anything Model (SAM) is proposed, demonstrating the effectiveness and generalizability of the proposed framework across diverse remote sensing s...

Jing Li, Yu-Lin Cao, Xian-Tao Jiang et al. · 0 citations
2026

Enhancing Scene Generalization for Open-Vocabulary Remote Sensing Segmentation via Semantic–Structural Collaboration

Open-vocabulary semantic segmentation (OVSS) of remote sensing faces severe performance degradation when encountering unseen scene distributions caused by geographic, sensor, and resolution variations. Existing vision–language approaches provide strong semantic priors but lack scene-invariant structural representations...

Wu-Biao Huang, Hu-Chen Li, Shuai Zhang et al. · 0 citations
Open access 2026

Dual-Level Prototype Alignment via Cross-Attention for Few-Shot Remote Sensing Semantic Segmentation

DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...

Mustafa Alawadi, M. Fateh · 0 citations
#artificial intelligence Preprint Sep 2026

SatOV: Restoring Spatial Priors for Training-Free Open-Vocabulary Segmentation in Remote Sensing Imagery

Open-vocabulary semantic segmentation (OVS) of remote sensing imagery is a challenging pixel-level task requiring strong generalization and adaptation to the spatial characteristics of remote sensing data. Although existing vision-language foundation models perform well in general domains, their image-level classificat...

Chang-Hao Zhao, Ling-Lin Zeng, Hai Liu · 0 citations
2026

Focused Adapter: Enhancing Fine-Grained Attention for Remote Sensing Image–Text Retrieval

The emergence of large-scale vision–language models (VLMs) has significantly advanced remote sensing image–text retrieval (RSITR) by providing powerful cross-modal semantic priors. However, when adapted to the remote sensing (RS) domain, these models struggle to capture fine-grained representations due to their inheren...

Wen-Liang Du, Xiao-Yu Xu, Jia-Qi Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.