Skip to content
Open access

A Novel Label-Free Approach for Post-Fire Environmental Assessment Based on Zero-Shot Segment Anything Model (SAM)

Jul 2026 · The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences · Vol XLIX-B3-2026, pp. 551-556 · 0 citations · 20 references

TL;DR

A zero-shot burned-area mapping framework based on the Segment Anything Model (SAM) and multispectral Sentinel-2 imagery and integrating index-based composites with SAM outputs significantly enhances the discrimination between burned and unburned surfaces by reducing boundary fragmentation and spectral confusion in heterogeneous landscapes.

Abstract

Abstract. Accurate and timely burned-area delineation is essential for quantifying wildfire impacts on ecosystem functioning, carbon dynamics, and post-fire recovery. Conventional pixel-based approaches remain sensitive to spectral ambiguity, topographic effects, and empirically defined thresholds, while recent deep learning models (e.g., U-Net, DeepLab, SegFormer) are constrained by their dependence on large, site-specific labelled datasets and repeated regional retraining. This study proposes a zero-shot burned-area mapping framework based on the Segment Anything Model (SAM) and multispectral Sentinel-2 imagery. Composite representations derived from ΔNBR, ΔNBR2, and ΔNDVI were generated and used as primary inputs to SAM in a label-free configuration. The effects of alternative pre-processing strategies, post-processing operations, and key hyperparameter settings were systematically investigated. Results show that multi-scale inference (crop_n_layers = 2) substantially improves geometric consistency and boundary accuracy of the extracted burned-area masks. The highest Intersection over Union values reached 0.89 for the Bursa study site and 0.87 for the Çanakkale study site, with corresponding F1 scores of 0.94 and 0.92, respectively. Despite the complete absence of training samples, SAM achieves performance comparable to, and in some cases exceeding, that of supervised deep learning approaches. Furthermore, integrating index-based composites with SAM outputs significantly enhances the discrimination between burned and unburned surfaces by reducing boundary fragmentation and spectral confusion in heterogeneous landscapes. By eliminating the need for manually labeled training data, the proposed framework addresses a major operational bottleneck in deep learning–based remote sensing. Overall, the study demonstrates a fast, scalable, and cost-effective solution for operational burned-area mapping and highlights the strong potential of SAM for zero-shot environmental monitoring and rapid post-fire response.

Read PDF

Similar papers

Preprint Sep 2026

Scale-based Approach for Active Wildfire Segmentation on Satellite Imagery

Active wildfire mapping from satellite imagery is challenging due to the sparse and highly imbalanced nature of fire pixels, especially in early-stage or low-density fire observations. This work investigates the use of multispectral Landsat-8 imagery for active-fire segmentation under multi-scale wildfire size conditions. We propose a data-driven protocol to characterize fire-region size distributions through connected-component analysis and an interquartile range criterion, enabling the evaluation of model robustness across different local fire-region densities. Three segmentation architectures, U-Net, DeepLabV3+, and SegFormer, are evaluated under different SWIR-based spectral configurations. Results show that U-Net achieves the strongest robustness across the evaluated conditions, SegFormer provides competitive performance, and DeepLabV3+ tends to produce conservative predictions with reduced recall. Across architectures, SWIR2 consistently achieves the strongest or near-best results, highlighting its importance for active-fire segmentation in Landsat-8 imagery. These findings suggest that both spectral band selection and architectural design are critical for robust satellite-based active wildfire mapping trained on low active fire-pixel density images.

Matheus F. Kovaleski, C. Premebida, J. Paulo · 0 citations
Open access 2026

Hyperspectral Knowledge-Guided Cross-Sensor Burned-Area Mapping: Distilling AVIRIS Semantics to Sentinel-2 and Landsat-8

Accurate burned-area mapping is essential for postfire assessment, vegetation recovery monitoring, carbon-emission estimation, and ecological-restoration planning. Multispectral satellites enable scalable burned-area monitoring, but limited spectral resolution can confound fire effects with agricultural disturbance, exposed soil, shadows, and other nonfire changes. This study proposes a task-oriented hyperspectral-to-multispectral semantic-distillation framework in which bitemporal airborne visible/infrared imaging spectrometer (AVIRIS) observations supervise deployable multispectral students. The evaluated real-target pathway maps AVIRIS samples to Landsat-8 or Sentinel-2 pixels and trains the student with blended hard labels and teacher probabilities. An optional spectral-response-function-only pathway is formulated for cases without paired target observations but is not quantitatively evaluated. Under grouped leave-one-scene-out evaluation on five Landsat-8 scenes, the strongest 101-seed pairing—a transformer teacher and a multilayer perceptron (MLP) student—achieved 89.17 ± 1.99% mean F1-score and 81.05 ± 2.92% burned-class intersection over union (IoU), improving direct MLP by 1.64 and 2.22 percentage points. The compact MLP-to-MLP pairing also improved its direct counterpart. On three Sentinel-2 scenes, this compact pairing improved direct MLP by 3.82 F1 points and 5.53 IoU points, whereas Mamba-style results showed architecture-dependent gains. Independent evaluation across Australia, Chile, and Mediterranean islands showed that the lightweight student achieved competitive precision–recall performance without regional retraining, with band-angle-index random forest and Rao’s Q as event-level comparators and the global annual burned-area map as an annual product-level reference. The results support AVIRIS-guided semantic transfer when paired training records are available, while broader validation, hard-scene stabilization, and dedicated evaluation of the optional simulated pathway remain necessary.

Chun-Xiu Liu, Lin Sun, Yong Chen · 0 citations
#machine learning Review Sep 2026

A Sentinel-2 benchmark dataset for deep-learning active-fire segmentation across 25 California wildfires

This article describes an open image dataset for developing and evaluating active-fire segmentation methods in satellite imagery. The dataset contains 2,148 image-mask pairs from 25 California wildfires, with acquisitions spanning July 2020 to August 2026. Each image is a 512x512-pixel, three-channel composite derived from Sentinel-2 Level-2A bands B12, B11 and B8A at 20 m spatial sampling. A fixed linear rendering is applied throughout the dataset. Corresponding masks distinguish background, SWIR-rule active fire and invalid observations. The masks were generated from shortwave-infrared brightness and near-infrared contrast, followed by constrained neighborhood growth. The release includes chip-level metadata and an incident-disjoint partition containing 18 training, three validation and four test fires. Among the image pairs, 841 contain active-fire labels; these labels occupy 0.0766% of all grid cells. A mask-blind analyst review covers 233 test chips and provides a separate assessment of the rule-generated labels at chip and connected-component levels. Reference training and evaluation code accompanies the data, including a ResNet-34 U-Net implementation with validation-based checkpoint and threshold selection. The archived images, masks, metadata and review annotations support research on rare-class segmentation, learning from algorithmic labels and transfer across fire incidents. The versioned dataset is deposited on Zenodo, with preparation and reuse software maintained in a public GitHub repository.

Shreyan Mitra, Mohammadreza Narimani, Parastoo Farajpoor · 0 citations
Open access Aug 2026

Hybrid ViT-UNet Framework for Accurate River Segmentation and Buffer Zone Mapping in High-Resolution Satellite Imagery

This paper introduces a deep learning model for effective segmentation of rivers, lakes, and reservoirs from high-resolution Gaofen-2 satellite images, and demonstrates the potential of transformer-based segmentation models for remote sensing achieved accuracy of 98% for environmental risk management and decision support in disaster-prone areas.

T. S. Murthy, K. Rao, Swathi Sowmya Bavirthi · 0 citations
Open access Aug 2026

SpaSE-UNet3D: Sensor-Driven Wildfire Detection and Progression Prediction from VIIRS Multispectral Imagery

Timely wildfire monitoring depends critically on optical and thermal infrared sensor observations from spaceborne instruments. The TS-SatFire benchmark (2025) consolidates multispectral VIIRS image stacks from Suomi-NPP and NOAA-20 for three tasks: active fire (AF) detection, burned area (BA) mapping, and fire progression (FP) prediction. We make two contributions. First, a systematic label-quality audit reveals that many fires lack ground-truth annotations; 18 training fires and 2 test fires were excluded for AF, and the two unannotated test fires cannot be scored by any model. We further document the benchmark’s scoring procedure, which differs from ours in ways that make the two sets of figures incomparable, and the BA label encoding in the released GeoTIFFs; the BA task is only audited. Second, we propose SpaSE-UNet3D, a spatial squeeze-and-excitation 3D U-Net whose spatial-only (1,3,3) convolutions avoid temporal mixing on short observation windows, while SE channel attention reweights the VIIRS spectral bands dynamically. With micro-averaging over all test pixels, it reaches F1 = 0.8549±0.0005 on AF and 0.3845±0.0221 on FP at TS = 2, matching or exceeding the strongest published baselines on their respective terms. A single-day AF input reaches 0.8520±0.0008, within 0.003 of the two-day figure, indicating that one acquisition carries most of the detectable signal, whereas published baselines use up to six days; on FP, we use one third of their temporal context. An ablation shows the spatial-only design matches the accuracy of a full (3,3,3) network with 2.72× fewer parameters. Code and results are publicly available.

Nikolaos Mavros, Dimitrios Katsaros · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.