Aug 2026· The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences· Vol L-4/W1-2026, pp. 95-102· 0 citations· 13 references
TL;DR
The combination of Embedding V1, MLP, and Aitchison distance loss achieved the best overall performance, suggesting that foundation model embeddings combined with compositional losses can improve sub-pixel land-cover estimation.
Abstract
Abstract. Accurate land-cover maps are essential, but medium-resolution imagery (e.g., Sentinel-2 at 10 m) often contains mixed pixels that include multiple land-cover types. Standard “hard” classification assigns one class per pixel, hiding minority classes and reducing map usefulness. Compositional classification instead estimates the proportion of each class within a pixel, preserving sub-pixel detail, but requires outputs that are non-negative and sum to one, constraints not naturally handled by typical ML/DL losses. This study proposed and evaluated a deep-learning framework for compositional land-cover estimation at 10 m resolution. It compared two input feature sets: (1) reflectance from 10 Sentinel-2 multispectral bands (B2–B8, B8A, B11, B12) and (2) Embedding V1, a 64- dimensional representation from the AlphaEarth Foundations model that integrates multi-source, multi-temporal Earth observation signals. Ground-truth composition vectors were derived from OpenEarthMap by aggregating 0.25–0.5 m labels to 10 m pixels for eight classes. Three architectures (MLP, 2D-CNN, 3D-CNN) used Softmax outputs to enforce the constant-sum constraint, and two losses (MAE vs Aitchison distance) were tested. Embedding V1 improved estimation accuracy across all model architectures compared to Sentinel-2 spectral bands alone. While 3D-CNN achieved the best performance with Sentinel-2 input (MAE: 0.1126), MLP outperformed all other architectures when Embedding V1 was used (MAE: 0.0989). Comparison of fraction maps revealed that MAE produced spatially smoothed outputs, whereas Aitchison distance yielded sharper and more realistic compositions. The combination of Embedding V1, MLP, and Aitchison distance loss achieved the best overall performance, suggesting that foundation model embeddings combined with compositional losses can improve sub-pixel land-cover estimation.
The proposed GeoRGMAE, a geospatially guided masked autoencoder pretraining strategy for building segmentation, introduces three masking strategies that prioritize semantically relevant building regions under the varying urban densities and suggests that incorporating geospatial priors into masked image modelling (MIM) can improve representation learning for downstream building segmentation tasks.
Tuğba Eraslanoğlu, G. Mutreja, Martin Kada et al.· The International Archives o...· 0 citations
Abstract. For an increasing number of applications, land cover maps can be generated from remote sensing imagery using conventional and deep-learning-based semantic segmentation models. Relying on a large pool of training data, the networks struggle with the spatial-temporal-spectral heterogeneity in the complex and diverse remote sensing imageries, leading to a significant number of errors in the model predictions. This paper presents a workflow comprising domain adaptation and classification. In particular, we analyze two domain adaptation techniques: First, a conventional histogram-matching method, which has turned out to be a surprisingly fast and reliable tool in a previous study, and second, a CycleGAN, which we applied both in its standard form and with the perceptual loss, thereby penalizing style inconsistencies on deeper layers. By applying the workflow to three remote sensing datasets and six directions of domain adaptation, we show that there is “no free lunch” in the sense that all domain adaptation methods have their advantages. Depending on the dataset, classification method, and especially on the availability of 3D data, the performance gap can be reduced to up to 1.5% of the mean F1 score, demonstrating the soundness of the proposed method.
Edwin Deisling, Raphael Zipperer, B. Kottler et al.· The International Archives o...· 0 citations
The results establish zero-shot NAS as a computationally efficient paradigm for large-scale Earth observation segmentation as a training-free strategy for semantic segmentation in Earth observation.
Gabriel Iuhasz, Marian Neagul· IEEE Access· 0 citations
The remote sensing scene classification (RSSC) task plays a pivotal role in Earth observation missions, yet its progress remains constrained by the scarcity of high-quality labeled imagery. This article introduces a self-supervised learning (SSL) paradigm to address this challenge. First, for pseudo-label construction, a large set of long-interval satellite revisit imagery is collected and processed with pixel-level registration. The SIFT inliers retained during registration serve as saliency priors to guide asymmetric masking across views. This produces positive pairs that preserve global scene consistency while introducing controlled object-level ambiguities. Second, we propose a progressive layer-wise contrastive learning framework (MTC-Net) that couples the pseudo-label with the network’s representational hierarchy, forming a curriculum from local texture robustness to global semantic invariance. A dual-attention module with spatial–channel branches is further embedded to recalibrate intermediate features. The learning paradigm encourages the model to perform cross-view contextual reasoning rather than relying on pixel-wise correspondences. Experiments on three widely used datasets demonstrate that MTC-Net achieves competitive classification accuracy under limited-label settings, while ablation and visualization studies validate the effectiveness of establishing scene-level invariance through multi-temporal contrastive alignment.
Xiao Xiao, Han Zhang, Kenan Cheng et al.· Remote Sensing· 0 citations