Jul 2026· IEEE Transactions on Geoscience and Remote Sensing· Vol 64, pp. 3002614-3002614· 0 citations· 74 references
Computer Science
TL;DR
Disaster TD, a disaster toponym disambiguation framework that integrates multimodal large language models (MLLMs)-based semantic reasoning with cross-view geolocalization, is proposed, demonstrating the effectiveness of integrating MLLM-based candidate generation with cross-view verification for fine-grained disaster geolocalization.
Abstract
Social media imagery (SMI) provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency response. Unlike satellite or aerial imagery, SMI can capture disaster impacts and ground-level conditions in a timely manner. However, geographic references in SMI are often vague or ambiguous, making accurate geolocalization challenging. To address this issue, we propose Disaster TD, a disaster toponym disambiguation framework that integrates multimodal large language models (MLLMs)-based semantic reasoning with cross-view geolocalization. First, MLLMs extract toponyms and generate candidate geolocations from noisy textual inputs. Then, cross-view matching between SMI, remote sensing imagery (RSI), and optionally street-view imagery (SVI) is used to verify and refine these candidate results. A Vision Transformer (ViT)-based visual foundation model, DINOv2, is used to bridge the domain gap between overhead and ground-level imagery. We evaluate DisasterTD on the Hurricane Harvey dataset, where SMI is augmented with collected RSI and SVI to construct a cross-view benchmark for disaster geolocalization. The dataset is divided into four categories based on toponym clarity and ambiguity, allowing a fine-grained performance analysis across scenarios. Results show that DisasterTD consistently outperforms MLLM-only and cross-view-only baselines without disambiguation, achieving geolocalization accuracies of 71.62% within 1000 m, 62.36% within 500 m, 57.99% within 250 m, 52.09% within 100 m, and 47.01% within 50 m, while reducing the mean and median errors to 11.33 and 0.68 km, respectively. The largest improvements appear in ambiguous toponyms, where semantic reasoning with cross-view evidence reduces candidate dispersion and errors. These findings demonstrate the effectiveness of integrating MLLM-based candidate generation with cross-view verification for fine-grained disaster geolocalization.
Abstract. Rapid identification of disaster locations is essential for effective emergency response and situational awareness. However, a large proportion of images shared on social media during disasters lack geographic metadata, limiting their usefulness for operational decision-making. Existing approaches mainly rely on geotags or textual geoparsing, which often fail when metadata is missing or ambiguous. As a result, valuable visual information from social media images remains underutilized for disaster mapping. This study proposes a deep learning–based framework to estimate the geographic location of disaster images using visual scene matching. The approach compares query images from social media with georeferenced Google Street View imagery to infer their potential locations. The framework integrates image pre-processing, deep feature extraction, and similarity-based matching to identify the most likely geographic correspondence. Preliminary experiments demonstrate promising results in detecting key scene elements within disaster imagery, providing a foundation for reliable visual matching. By leveraging visual cues rather than textual metadata, the proposed framework aims to improve the usability of non-geotagged social media images for disaster response. The proposed approach has the potential to support rapid disaster mapping by transforming citizen-generated imagery into spatially actionable information for emergency management.
Pai-Hui Hsu, M. Sathianarayanan· The International Archives o...· 0 citations
DamageScope is a retrieval-augmented framework that combines satellite imagery with Vision-Language Models (VLMs) and Large Language Models (LLMs) to automate property damage analysis and introduces a novel multi-vector embedding-based clustering algorithm that outperforms traditional single-vector embedding approaches while reducing indexing time by up to 14x.
Ravi K. Rajendran, Biplob Debnath, Murugan Sankaradas et al.· 0 citations
Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite recent advances, current methods struggle with cross-region generalization and semantic interpretability due to their reliance on region-specific auxiliary data and the neglect of semantic alignment within multi-temporal urban imagery. Therefore, we present CoST, a novel \underline{Co}ntrastive-based \underline{S}patial-\underline{T}emporal framework that aligns spatial context with multi-temporal semantics to extract universal geographic regularities shared across regions. Specifically, CoST explicitly models spatial correlations to capture transferable geographic structures and exploits multi-year urban change semantics to align learned representations with high-level geo-semantics. Extensive experiments demonstrate that CoST consistently achieves superior performance across various downstream tasks and in unseen scenario, yielding an average relative gain of 8.7\% over the strongest competing methods across eight city-indicator settings. The code is available in \href{https://github.com/Arandinglv/CoST}{this repo}.
Yutian Jiang, Jiabo Liu, Xixuan Hao et al.· 1 citation
Remote-sensing multimodal large language models (MLLMs) often assert facts that imagery cannot establish, such as a facility's identity or function. Coordinate-keyed geographic retrieval can supply this missing knowledge, improving fMoW land-use accuracy by 12.06--17.19 points across three open MLLMs. However, retrieved records can also contradict visible evidence, and we find that models frequently follow the records even when the image is decisive. We argue that source trust should therefore depend on \emph{cross-modal verifiability}: geographic records are most useful for attributes the image cannot verify and most dangerous when they dispute visually verifiable attributes. We introduce GeoArbiter, a training-free pipeline that operationalizes this principle by injecting only image-unverifiable geographic facts. Unlike arbitration prompts, which leak across attribute types and bias yes/no responses, content-level filtering preserves 84.69--87.15\% of the full-retrieval accuracy gain, reduces claim-level hallucination by 9.58--26.34\% under a source-blinded judge, and improves robustness to conflicting records across all three models. These results identify verifiability-guided content selection as a simple, effective mechanism for grounding remote-sensing MLLMs in fallible geographic knowledge.
Street-scale telecom growth requires jointly reasoning over heterogeneous evidence, including geospatial context (buildings, communities, and business districts), network capability (coverage and capacity), and customer-side signals (service usage and support tickets). Existing LLM-based sales assistants are often ungrounded: they ignore deliverability constraints, lack verifiable evidence, and fail to produce actionable plans that can be executed by field teams. In this paper, we introduce StreetCopilot, a grounded vision-language agent for street-level prospecting and service planning. StreetCopilot integrates (i) geospatial visual cues (e.g., street-view/remote-sensing building context and POIs), (ii) structured telecom signals (coverage maps, traffic KPIs, product portfolios), and (iii) unstructured operational text (tickets and visit notes) via a retrieval-augmented reasoning pipeline. To ensure actionability, we propose a deliverability-aware constraint module that verifies whether recommended bundles (network, wireless coverage, industrial devices, and cloud services) are feasible under local coverage and resource conditions, and a citation-grounded generation mechanism that attaches evidence snippets to each recommendation for auditability. We further present a street-scale closed-loop evaluation protocol that measures not only recommendation accuracy but also plan feasibility, evidence faithfulness, and end-to-end business outcomes (lead acceptance and conversion). Experiments on a real-world deployment in a city subregion demonstrate that StreetCopilot substantially improves prospect ranking quality and proposal drafting efficiency while maintaining high feasibility and evidence faithfulness, shedding light on grounded multimodal agents for real-world decision-making.
junchi ren· Poster Volume 0008 The 2026...· 0 citations
Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving unclear whether unified embedding space can support heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal changes. To address this gap, we make three key contributions. First, we introduce GeoMEB, a large-scale multimodal embedding benchmark that standardizes 45 urban evaluation tasks across retrieval, visual question answering, change detection, classification, and visual grounding, together with training collections comprising 1.32M examples and 286K evaluation queries. Second, we present Geo-Embed, a unified embedding model that adapts a shared vision-language backbone to instruction-conditioned query-target matching over heterogeneous geospatial inputs, including single images, multiple images, text, regions, and masks. On GeoMEB, Geo-Embed achieves the strongest overall performance among representative multimodal embedders, with a 15.3% relative improvement over the strongest baseline. These results motivate future geospatial embedders that organize training and evaluation around explicit query-target relations, including semantic, cross-view, region-level, and temporal correspondence.
Jia-Peng Li, Yong Li, Junjie Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.