Multiclass Cloud Detection for Very High-Resolution Satellite Imagery Using U-Net: A Multi-Sensor Validation Study
Abstract
Cloud detection is a critical preprocessing step in very high-resolution (VHR) optical satellite image analysis, yet existing methods predominantly address binary cloud/non-cloud classification and are limited to single-sensor, single-region evaluations. This paper proposes a deep learning approach for multiclass cloud detection—thick clouds, thin clouds, cloud shadows, and clear areas—in WorldView-3 imagery using U-Net, with an additional cross-sensor evaluation on SPOT 6/7 imagery. Using all four available spectral bands (blue, green, red, near-infrared), we apply a percentile-based contrast enhancement strategy to compensate for the absence of thermal infrared bands in VHR sensors. Trained on 20 WorldView-3 tiles from Indonesian tropical regions (Flores) with on-the-fly augmentation and evaluated on a 3-tile held-out test set, the proposed U-Net achieves mIoU<inline-formula> <tex-math notation="LaTeX">$ = 0.7545$ </tex-math></inline-formula> (0.7618 with test-time augmentation), mean F<inline-formula> <tex-math notation="LaTeX">$1 = 0.8586$ </tex-math></inline-formula>, and <inline-formula> <tex-math notation="LaTeX">$\kappa =0.8097$ </tex-math></inline-formula>; spatial five-fold cross-validation on the full 25-tile dataset yields mean mIoU <inline-formula> <tex-math notation="LaTeX">$=0.7229\pm 0.0732$ </tex-math></inline-formula>. Cross-sensor evaluation on genuine SPOT 6/7 PMS imagery over Yogyakarta shows substantial zero-shot degradation (mIoU from 0.7254 to 0.0648), improving to 0.5109 after few-shot fine-tuning; a controlled comparison against raw-input and no-augmentation variants shows neither design rationale is supported under this sensor shift. Ablation studies identify Batch Normalisation removal as causing severe class collapse (mIoU<inline-formula> <tex-math notation="LaTeX">$ = 0.364$ </tex-math></inline-formula>) and NIR band removal as reducing cloud shadow F1 by 0.326. Comparative experiments against six baseline architectures confirm competitive performance, and failure case analysis identifies terrain shadow confusion and thin cloud over bright surfaces as the primary remaining challenges.