RAH-VLA: Resolution-Adaptive Hierarchical Vision–Language Alignment for Multimodal Remote Sensing Understanding
Abstract
Multimodal vision–language modeling has emerged as a promising paradigm for remote sensing (RS) image understanding. However, existing methods are limited by fixed-resolution visual processing and single-scale vision–language alignment, making it difficult to simultaneously preserve fine-grained details and maintain semantic consistency across different spatial granularities. To address these challenges, we propose RAH-VLA, a resolution-adaptive hierarchical vision–language alignment framework for multimodal RS understanding. Specifically, a dynamic resolution input strategy (DRIS) is developed to enable resolution-adaptive visual representations, while a multiscale vision–language alignment mechanism (MS-VLAM) is introduced to establish hierarchical semantic correspondence across object-level, region-level, and global-level representations. Extensive experiments on multiple RS benchmarks demonstrate that RAH-VLA consistently improves image captioning, visual grounding, and cross-modal reasoning performance while reducing computational redundancy. Qualitative analyses further illustrate the effectiveness of the proposed resolution-adaptive perception and hierarchical vision–language alignment mechanisms, as well as the cross-modality generalization capability of the proposed framework. Overall, RAH-VLA provides an effective and scalable solution for multimodal RS interpretation, offering a practical pathway toward efficient and semantically robust RS vision–language models.