Parameter-Efficient Local-Context Cooperation via Vision Foundation Models for UHR Remote Sensing Image Segmentation
Abstract
Ultrahigh resolution (UHR) remote sensing image segmentation aims to achieve a fine-grained understanding of complex ground scenes. In recent years, vision foundation models (VFMs) have shown strong capability in learning generic structural priors from large-scale visual data, indicating great potential for such fine-grained scene understanding. However, their application to UHR remote sensing images remains limited, as the massive parameter scales of VFMs are difficult to train under UHR remote sensing images. Motivated by the success of the parameter-efficient fine-tuning paradigm on VFMs, we propose a novel parameter-efficient local-context cooperation (PEACE) framework, which significantly reduces trainable parameter overhead while improving segmentation accuracy. In particular, PEACE leverages a shared VFM with minimal trainable parameters to collaboratively process local and corresponding contextual patches partitioned from the UHR remote sensing image. A multireceptive local adapter (MRLA) and a multireceptive context adapter (MRCA) are designed to capture spatial features of local and contextual inputs across multiple receptive fields. Finally, contextual semantics are integrated into local representations. Furthermore, a context-sensitive assistance strategy (CSAS) leverages the correct prediction of the context to effectively overcome the primary limitations of the patch-based training paradigm. Experimental results demonstrate that PEACE effectively exploits VFMs and remote sensing foundation models (RSFMs) with minimal parameter increments and achieves versatility and superior performance across several UHR remote sensing image benchmarks.