Skip to content

Author

Wen-Liang Du

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

2026

Focused Adapter: Enhancing Fine-Grained Attention for Remote Sensing Image–Text Retrieval

The emergence of large-scale vision–language models (VLMs) has significantly advanced remote sensing image–text retrieval (RSITR) by providing powerful cross-modal semantic priors. However, when adapted to the remote sensing (RS) domain, these models struggle to capture fine-grained representations due to their inherent global semantic bias. To address these limitations, we propose the focused adapter (FocA), a plug-and-play parameter-efficient fine-tuning (PEFT) architecture designed to enhance fine-grained perception from frozen VLMs. The FocA features a hybrid structure consisting of two components: an enhancement adapter and an alignment adapter. The enhancement adapter utilizes bottleneck self-attention to capture rich patch-level details often lost during global pooling. The alignment adapter incorporates a cross-modal shared projection subspace to facilitate early feature interaction and implicit alignment. In addition, we develop an explicit shared loss to provide direct semantic supervision, preventing the dilution of critical features during deep propagation. Extensive experiments on the RSICD and RSITMD datasets demonstrate that FocA achieves state-of-the-art (SOTA) performance, attaining a mean recall (mR) of 37.61% on RSICD and 48.95% on RSITMD. Our code is available at https://github.com/WenliangDu/FocA

Wen-Liang Du, Xiao-Yu Xu, Jiaqi Zhao et al. · 0 citations
2026

Toward Zero-Forgetting: A Training-Free Multimodal Framework for Remote Sensing Class-Incremental Learning

Existing class-incremental learning (CIL) methods for remote sensing (RS) scene classification often tend to be training-intensive or rely on static visual features that may inadequately capture the complex interclass similarity and intraclass diversity inherent in RS imagery. Moreover, directly reusing features from models pretrained on natural images (e.g., CLIP) often suffers from a domain gap and limited semantic richness when transferred to the RS domain. In this article, we propose a simple yet effective training-free CIL framework for RS scene classification that leverages multimodal semantic information to build more discriminative category representations. Our framework treats pretrained models as frozen feature extractors to guarantee zero forgetting of the representation space. To enhance semantic discriminability, we employ a large-language model (LLM) to generate rich candidate textual descriptions for each class and introduce an image-guided description selection (IGDS) strategy to align semantic information with visual characteristics. Classification is performed using a distance-based metric without any additional training. Extensive experimental results demonstrate that our framework achieves leading performance and superior stability across different session sequences, surpassing both training-based and training-free baseline methods. Our code is available at https://github.com/WenliangDu/ZFCIL-RS

Wen-Liang Du, Ji-Cun He, Jiaqi Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.