Skip to content

MFS: a saliency-driven interactive multimodal fusion framework for robust semantic segmentation and object detection in complex and occluded scenes

Aug 2026 · Multimedia Systems · Vol 32 · 0 citations · 81 references

TL;DR

An interactive multimodal scene understanding framework based on frequency-domain dynamic routing and activation-region guidance, aiming to enhance multimodal feature representation for semantic segmentation and object detection, consistently outperforms existing approaches in image fusion, semantic segmentation, and object detection.

View source

Similar papers

Open access Aug 2026

STSFusion: segmentation task-driven spatial-frequency collaborative fusion for infrared and visible images

Infrared and visible image fusion aims to simultaneously preserve the saliency of thermal targets and rich texture details. Most existing methods primarily rely on spatial-domain representations, while the frequency-domain information is not sufficiently explored. Moreover, it remains challenging to simultaneously main...

Yutong Chen, Zhe Hu, Yang Li et al. · 0 citations
Aug 2026

MSPD-net: structural–appearance prototype decoupling for weakly supervised semantic segmentation

A complementary prototype representation framework is proposed, employing three modules to collaboratively improve pseudo-label quality and improves the discriminative ability of confused categories by generating semantically similar sub-category negative samples.

Wei Cao, Yong Jiang, Ruiying Wang · 0 citations
Jul 2026

TFNet: a triple-fusion network for multispectral pedestrian detection

Abstract. Occlusions and far distance pedestrians pose significant challenges for pedestrian detection, often leading to insufficient feature representation learned by models, which in turn results in degraded detection accuracy and a high miss rate. To address this issue, we propose a three-stage fusion multispectral...

Chao-Wen Chai, Yan-Ni Wang, Xiang Hu et al. · 0 citations
Open access Aug 2026

MAEF-Net: An Efficient Multi-Scale Attention-Enhanced Feature Fusion Network for Remote Sensing Object Detection

(1) Objective: Remote sensing object detection faces significant challenges, including complex background interference, large variations in target scales, and insufficient multi-scale feature representation, which often result in missed detections of small objects, inaccurate localization, and inadequate feature fusion...

Hongyan Shi, Xiaofeng Bai, Chenshuai Bai · 0 citations
Open access Sep 2026

A small and dense target detection framework based on feature perception and adaptive fusion

Small and dense object detection remains challenging in complex visual scenes. Repeated downsampling weakens discriminative features of tiny objects, while dense object distributions cause severe feature overlap and semantic ambiguity. To address these challenges, this paper proposes Enhanced Feature-Aware YOLO (EFA-YO...

Zhen Zhang, Xu Xie, Yi Zhang et al. · 0 citations
2026

MAMENet: Modal Alignment and Multiscale Feature Enhancement Network for RGB–IR Small Object Detection

Object detection using visible–infrared images has become increasingly important for all-day detection scenarios. However, due to significant imaging discrepancies between the visible and infrared modalities, achieving accurate modal alignment and effective feature fusion remains a major challenge. Existing methods oft...

Kai-Yue Men, Cheng-You Wang, Xiao Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.