CRISP: Cross-Modal Residual Guidance and Spatial Realignment for Remote Sensing Visual Question Answering
Remote sensing visual question answering (RSVQA) remains challenging because the evidence relevant to the question in remote sensing (RS) imagery is typically sparse, spatially dispersed, and highly variable in scale and layout. This structural heterogeneity poses a major challenge to existing transfer strategies, whic...