MFS: a saliency-driven interactive multimodal fusion framework for robust semantic segmentation and object detection in complex and occluded scenes
An interactive multimodal scene understanding framework based on frequency-domain dynamic routing and activation-region guidance, aiming to enhance multimodal feature representation for semantic segmentation and object detection, consistently outperforms existing approaches in image fusion, semantic segmentation, and o...