Skip to content

Dual-Prompt-Driven Cross-Modal Fusion Learning for Few-Shot Hyperspectral Image Classification

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 4417117-4417117 · 0 citations · 42 references

Abstract

Vision-language multimodal learning has exhibited remarkable advantages in few-shot hyperspectral image (HSI) classification, where prompt learning effectively enhances feature-extraction accuracy and representation quality by guiding the model to focus on critical information. However, static prompts lack flexibility, while dynamic prompts suffer from unstable generation. To address these issues, this article proposes a dual-prompt-driven cross-modal fusion learning (DPCFL) method. Specifically, we utilize static prompt templates to provide stable prior guidance for the model. Simultaneously, a unique learnable prompt vector is designed for each category, which operates independently of the pretrained model’s input to effectively mitigate potential influences from prior semantics. To alleviate semantic bias, we further design a multiloss joint optimization strategy incorporating parameter space, feature space, and cross-modal space consistency constraints, thereby improving the robustness of class-level prototype features. In addition, to tackle data scarcity, a cross-domain collaborative training mechanism is introduced to facilitate knowledge transfer. The experimental results on multiple standard HSI datasets confirm its superior classification performance under both few-shot and cross-domain scenarios. The related code will be made publicly available at the following URL: https://github.com/AIYAU/DPCFL

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.