Skip to content

Multi-modal interaction enhanced segment anything model (MIE-SAM) for RGB-T salient object detection.

Sep 2026 · Neural Networks · Vol 205 Pt C, pp. 109586 · 0 citations · 53 references
Medicine

TL;DR

A Multi-modal Interaction Enhanced Segment Anything Model (MIE-SAM) that reconfigures SAM's image encoder into a weight-sharing dual-branch image encoders, and translates the fused features into the fine-grained saliency map in an entirely prompt-free, end-to-end manner.

Abstract

RGB-T Salient Object Detection (RGB-T SOD) effectively leverages the complementary information of RGB images and thermal infrared images to locate important targets in complex environments, such as low light, rainy and foggy weather, or cluttered backgrounds. However, the existing deep-learning based models have two key problems: the fixed fusion strategy can not adapt to the changing environmental conditions, and the scarcity of pixel-level annotation leads to over-fitting. While the Segment Anything Model (SAM) exhibits excellent generalization, its adaptation to RGB-T SOD is challenged by the lack of saliency semantics, RGB-only pre-training and manual prompt dependency. To address these challenges, we propose a Multi-modal Interaction Enhanced Segment Anything Model (MIE-SAM). The framework reconfigures SAM's image encoder into a weight-sharing dual-branch image encoders. A Multi-modal Low-Rank Adaptation module (Mm-LoRA) is embedded in the frozen dual-branch image encoders to inject saliency-specific semantics and facilitate deep cross-modal feature interaction while preserving SAM's pre-trained knowledge via parameter-efficient fine-tuning. Furthermore, a Dynamic Fusion Module (DFM) learns dynamic fusion weights to adaptively aggregate multi-modal embeddings based on their environmental reliability, ensuring robust integration in varying environments. Finally, a Progressive Decoder Module (PDM) directly translates the fused features into the fine-grained saliency map in an entirely prompt-free, end-to-end manner. Extensive experiments on several publicly available RGB-T SOD datasets show that our method achieves state-of-the-art performance, exhibiting robustness and generalization ability in a variety of challenging scenarios. The code is available at https://github.com/Aazzz66/MIESAM.

View source

Similar papers

Aug 2026

Diff-MM: Exploring Pre-Trained Text-to-Image Generation Models for Unified Multi-Modal Object Tracking.

A unified multi-modal tracker Diff-MM is proposed by exploiting the multi-modal understanding capability of the pre-trained text-to-image generation model by harnessing the extensive prior knowledge in the generation model to achieve a unified tracker with uniform parameters for RGB-N/D/T/E tracking.

Shiyu Xuan, Ze-Chao Li, Jin-Hui Tang et al. · 0 citations
Open access Sep 2026

LDANet: a lightweight depth-aware framework for RGB-D salient object detection

As a foundational research in computer vision, salient object detection (SOD) has received widespread attention from researchers. However, existing methods still have notable limitations, mainly reflected in two aspects. (1) Crude multi-modal fusion strategies fail to emphasize consistent multi-modal features and prese...

Bao-Yu Wang, Ping-Ping Cao, Xiao-Chun Guo et al. · 0 citations
Open access Aug 2026

Multi-Modal RGB–Depth Image Segmentation Using Feature Fusion

Image segmentation remains a challenging task, particularly in complex environments where visual information from RGB images alone is often insufficient. Factors such as poor lighting, occlusions, and background clutter can significantly degrade segmentation performance. To address these limitations, multi-modal approa...

Noor Safa, Zainab Majeed Abid · 0 citations
Open access 2026

TFE-Fusion: Tri-Modality Feature Enhanced Fusion for Robust Object Detection

Robust and reliable object detection under adverse conditions remains a critical challenge for automated driving systems (ADS). The performance of RGB-based (visible-spectrum) cameras degrades in poor lighting conditions, whereas the performance of RGB-Long-Wave Infrared (LWIR) fusion architectures remains limited in a...

Muhammad Usama, Imad Ali Shah, Roshan George et al. · 0 citations
Preprint Sep 2026

RA-SOD: Reliability-Aware RGB-T Salient Object Detection under Modality Degradation

RGB-Thermal (RGB-T) salient object detection leverages complementary cues from visible and thermal modalities to improve robustness in challenging environments. However, in real-world scenarios, the reliability of each modality is inherently unstable: RGB images degrade under low illumination, motion blur, and noise, w...

Hong-Bo Gao, Zheng-Yu Li, Xue-Ru Nie et al. · 2 citations · ⚡1
Sep 2026

Phase Consistency Prior Driven RGB-D Salient Object Detection.

For RGB-D salient object detection (SOD), a fundamental challenge lies in establishing effective cross-modality interactions between the input graphic domain (RGB and depth modalities) and the output saliency domain. While existing deep learning methods primarily focus on modeling image-level consistency through carefu...

Jing-Yi Xu, Xin Deng, Minglang Qiao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.