Skip to content

RAH-VLA: Resolution-Adaptive Hierarchical Vision–Language Alignment for Multimodal Remote Sensing Understanding

Dec 2025 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 5637018-5637018 · 1 citation · 68 references
Computer Science

Abstract

Multimodal vision–language modeling has emerged as a promising paradigm for remote sensing (RS) image understanding. However, existing methods are limited by fixed-resolution visual processing and single-scale vision–language alignment, making it difficult to simultaneously preserve fine-grained details and maintain semantic consistency across different spatial granularities. To address these challenges, we propose RAH-VLA, a resolution-adaptive hierarchical vision–language alignment framework for multimodal RS understanding. Specifically, a dynamic resolution input strategy (DRIS) is developed to enable resolution-adaptive visual representations, while a multiscale vision–language alignment mechanism (MS-VLAM) is introduced to establish hierarchical semantic correspondence across object-level, region-level, and global-level representations. Extensive experiments on multiple RS benchmarks demonstrate that RAH-VLA consistently improves image captioning, visual grounding, and cross-modal reasoning performance while reducing computational redundancy. Qualitative analyses further illustrate the effectiveness of the proposed resolution-adaptive perception and hierarchical vision–language alignment mechanisms, as well as the cross-modality generalization capability of the proposed framework. Overall, RAH-VLA provides an effective and scalable solution for multimodal RS interpretation, offering a practical pathway toward efficient and semantically robust RS vision–language models.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.