RAH-VLA: Resolution-Adaptive Hierarchical Vision–Language Alignment for Multimodal Remote Sensing Understanding
Multimodal vision–language modeling has emerged as a promising paradigm for remote sensing (RS) image understanding. However, existing methods are limited by fixed-resolution visual processing and single-scale vision–language alignment, making it difficult to simultaneously preserve fine-grained details and maintain se...