VFM-MoME: A Remote Sensing Landslide Image Segmentation Network Guided by a Visual Foundation Model and a Mixture of Mamba Experts
The linear computational complexity embraced by Mamba has demonstrated significant application potential in context modeling for the landslide segmentation tasks from remote sensing images. However, existing methods show deficiencies in terms of discrimination and generalization when applied to extreme remote sensing landslide scenarios, such as low resolution and abnormal lighting. To confront these challenges, we propose a remote sensing image landslide segmentation network (VFM-MoME) jointly guided by a vision foundation model and a mixture of Mamba experts. Specifically, we first design a dual-branch joint encoding architecture that integrates a frequency-aware wavelet block as the main encoding branch with the visual foundation model fusion as the auxiliary branch, thereby mitigating the issue of insufficient generalized features in specific landslide study areas. We also construct a mixture of Mamba expert block to enable the decoder to process both global context and local fine-grained features of landslides, addressing the shortcoming of simple serial Mamba in capturing local details and balancing between global semantic relationships and the edges and textural details of objects. Furthermore, we bring in a binary uncertainty enhancement module to guide the model in exploring challenging samples, thus enhancing the model’s ability to handle ambiguous features. Test results on the publicly available datasets of Landslide4Sense and GVLM demonstrate that our method achieves competitive performance.