Skip to content

ArtFusion: An Arbitrary-Reference Spatiotemporal Fusion Model for Seamless Remote Sensing Image Reconstruction

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 5632520-5632520 · 0 citations · 54 references

Abstract

Spatiotemporal fusion (STF) represents a vital solution for continuous high-resolution Earth observation. Despite the increasing abundance of historical satellite archives, most existing STF methods are constrained by a limited number of auxiliary reference images (typically one or two), leading to underutilized multitemporal information and reduced reconstruction reliability, particularly when References are compromised by cloud contamination or abrupt land-surface variations. To address these limitations, this study proposes an arbitrary-reference STF (ArtFusion) model. The architecture incorporates an efficient contrast-aware hybrid block (ECHB) for deep feature extraction, an explicit temporal information encoder (ETIE) to utilize acquisition metadata, and a multihead cross-reference attention fusion (MCAF) module designed to facilitate the integration of an arbitrary number of reference images. A comprehensive evaluation across 24 experimental cases in three representative study regions demonstrates that ArtFusion consistently outperforms four state-of-the-art (SOTA) benchmarks [multilevel feature fusion with generative adversarial network (MLFF-GAN), C-ROBOT, RealFusion, and frequency-selected differential fusion transformer (FSDFormer)]. Compared with the best-performing baseline among the four competing methods, ArtFusion achieves average improvements of 6.8% in spectral accuracy [root-mean-square error (RMSE)] and 15.3% in spatial accuracy [improved edge difference metric (iEDGE)]. Notably, high reliability is maintained across four challenging scenarios: rapid phenological changes, drastic morphological variations, highly heterogeneous landscapes, and frequent cloud contamination, while also demonstrating strong cross-regional transferability. Despite its superior performance, ArtFusion has an extremely compact structure with 0.19 million trainable parameters, only 2.2% of MLFF-GAN’s parameters. This work demonstrates the potential of leveraging multiple reference images to push the boundaries of STF accuracy, rather than merely increasing model complexity. This flexible multireference fusion scheme provides a promising pathway for robust, large-scale Earth observation in cloudy and dynamically changing landscapes. The source code is available at: https://github.com/Andy-cumt/ArtFusion-STF

View source

Similar papers

Open access Jul 2026

A Dynamically Weighted Framework for Adaptive Reference-Based Super-Resolution

Abstract. Satellite remote sensing is inherently constrained by a trade-off between spatial and temporal resolution. As a result, high-temporal-frequency sensors such as Geostationary Ocean Color Imager-II provide operationally valuable observations but at coarse spatial resolution. Reference-Based Super-Resolution (Re...

Chae-Eun Kim, Junhwa Chi · 0 citations
Conference Sep 2026

Beyond the Clouds: Reliable and Cloud-Aware Spatiotemporal Fusion via Adversarial Regression Wavelets

The Cloud-Aware Wavelet Generative Adversarial Network (CLAW-GAN), a novel framework for high-fidelity reconstruction under cloud-contaminated conditions, achieves state-of-the-art performance and demonstrates superior robustness across varying cloud coverage.

Si-Chen Lu, Ming-Fei Li, Juan-Juan Jing et al. · 0 citations
Open access Aug 2026

FlowT-SR: A Novel Remote Sensing Image Super-Resolution Framework with Cloud Haze and Noise Suppression

A novel SR framework based on the flow matching paradigm and a diffusion transformer, named FlowT-SR, which achieves superior and reliable reconstruction quality by jointly mitigating sensor noise and thin cloud interference, achieving superior reconstruction performance compared with current state-of-the-art methods i...

Yu-Tong Zhang, Guang Yang, Rong Liu et al. · 0 citations
Open access Aug 2026

Occlusion Removal in Remote Sensing Images Based on Deep Matrix Completion

Experimental results demonstrate that the proposed method consistently outperforms conventional matrix completion methods and achieves competitive performance compared with recent deep learning approaches, particularly under random missing patterns and high missing-rate scenarios.

Jie He, Zijian Lin, Tian-Yao Huang et al. · 0 citations
Jul 2026

Tri-scanning state-space model with multi-expert modulation for remote sensing image dehazing

TEMamba is presented, a tri-scanning state-space model with multi-expert modulation for remote sensing image dehazing, which converts feature representations into complementary scanning sequences along horizontal, vertical, and channel-related directions and achieves competitive restoration performance compared with ex...

Jun-Jie Li, Xin He, Yuan Feng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.