Skip to content
Preprint

GeoCore-9B: Towards Geo-Aware Generative Foundation Models in Earth Observation

Aug 2026 · 0 citations · 56 references
Computer Science

TL;DR

GeoCore-9B, a 9-billion-parameter generative foundation model, is introduced, which is the first of its scale to be trained from scratch exclusively on EO data and establishes new state-of-the-art performance in both visual fidelity and geographic structural accuracy.

Abstract

Existing generative models for earth observation (EO) predominantly rely on fine-tuning natural image priors, which limits their scalability and introduces perspective biases that conflict with geospatial constraints. To address this, we introduce GeoCore-9B, a 9-billion-parameter generative foundation model, which is the first of its scale to be trained from scratch exclusively on EO data. Unlike previous EO foundation models, GeoCore-9B is built upon a Flow Matching-based Diffusion Transformer (DiT) and natively conditions generation on text descriptions and continuous geospatial metadata, including ground sample distances, latitudes, and longitudes. To overcome the convergence and spatial disorientation challenges of training at this scale, we propose a Geospatial Semantic Alignment loss. This objective distills structural Earth surface priors (e.g., terrain and urban areas) from a frozen specialist teacher network, constraining the diffusion latent trajectory during training without adding inference overhead. Pre-trained on the global-scale Git-10M dataset, GeoCore-9B demonstrates strong downstream versatility. Beyond standard proxy generative tasks, we show that GeoCore-9B can be effectively adapted for practical EO applications, including highly challenging tasks such as cloud removal and SAR-to-optical cross-modal translation. Extensive evaluations confirm that GeoCore-9B establishes new state-of-the-art performance in both visual fidelity and geographic structural accuracy.

View source

Similar papers

Review

Earth Embedding Products for Geospatial Analysis: Foundations, Applications, and Open Challenges

This survey reviews the development of Earth embeddings from foundation model pretraining and reusable encoders to global and near-global embedding products, and summarizes the technical and scientific challenges surrounding global embedding products.

Yongchuan Cui, Ke-Li Shi, Fellow Ieee Shunlin Liang · 0 citations
Preprint Aug 2026

GeoBridge: Decoupled Semantic Conditioning for Generative Image Geolocalization

Multimodal large language models (MLLMs) have advanced image geolocalization mainly by improving how they reason about geographic cues. How that reasoning isdecoded into coordinates, however, has lagged behind. Predicting a place name for a geocoding API is discrete and lossy: it ignores image evidence and collapses mu...

Zhi-Yang Dou, Xumeng Han, Feng-De Peng et al. · 0 citations
Preprint Aug 2026

GeoPhysAdapter: Scale-Matched Geophysical Adaptation for Cross-Domain Landslide Mapping with Vision Foundation Models

Newly triggered landslides rarely carry immediate annotations, so cross-domain transferability determines the value of landslide mapping for emergency response and regional risk assessment. Vision foundation models have strengthened representational transfer, yet on unseen regions, events, and data sources they still g...

Zhi-Hang Liu, Mei-Po Kwan, Jin-Lin Wu et al. · 0 citations
Preprint Sep 2026

GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation

Paired synthetic aperture radar (SAR) and electro-optical (EO) imagery is increasingly available across sensors, resolutions, and geographic regions. Yet existing SAR-to-EO image translation (SET) methods are typically trained on a single, limited-scale dataset, producing models specialized to particular sensing condit...

Jeonghyeok Do, Munchurl Kim · 0 citations
#artificial intelligence Preprint Aug 2026

A Composition-Aware Pretraining Framework for Geospatial Foundation Models

Experimental evaluation shows that composition-aware pretraining yields substantial gains on region-level understanding tasks requiring semantic similarity judgment, including zero-shot image retrieval and scene classification, while remaining competitive on tasks requiring fine-grained spatial precision, such as segme...

A. Naveen, Abhishek Srinivas, Pranav Moothedath et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.