Skip to content

Dual-domain cross-attention fusion with edge-guided frequency decoupling for RGB-D saliency object detection

Jul 2026 · The Visual Computer · Vol 42 · 0 citations · 56 references
Computer Science

TL;DR

A Dual-Domain Confidence-Gated Cross Attention module (DD-CGCA), which adaptively fuses spectral and token-domain responses according to content-aware confidence-gated maps and outperforms the state-of-the-art SOTA models qualitatively and quantitatively due to multi-domain context information.

View source

Similar papers

Open access Sep 2026

Boundary-Guided Dual-Perspective Cross-Modal Fusion Network for RGB-IR Object Detection

A Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) is proposed to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives.

Hu Lin, Zhi-Wei Fu, Xiu-Mei Chen et al. · 0 citations
Open access Aug 2026

Structure-Prior-Guided Multi-Stage Cross-Modal Collaborative Network for RGB-D Semantic Segmentation

Red–green–blue and depth (RGB-D) semantic segmentation combines appearance cues from RGB images with geometric information from depth maps, but sensor noise, missing measurements, and boundary-inconsistent depth responses can introduce conflicting evidence during cross-modal fusion. We propose the Structure-Prior-Guide...

Yi-Fan Yu, Zhiwei Zhong, Fan Min et al. · 0 citations
Conference Aug 2026

DSDFNet: Dual-Stream Dense-Feedback Network for RGB-T Semantic Segmentation

RGB-T multi-modal semantic segmentation has demonstrated immense potential in complex scene understanding. However, existing methods are prone to introducing background noise during cross-modal feature interaction. Furthermore, traditional cascaded decoders frequently suffer from the dilution of deep semantic informati...

Hao-Yang Zhang, Xiao-Hong Dong, Nan Du · 0 citations
Sep 2026

Phase Consistency Prior Driven RGB-D Salient Object Detection.

For RGB-D salient object detection (SOD), a fundamental challenge lies in establishing effective cross-modality interactions between the input graphic domain (RGB and depth modalities) and the output saliency domain. While existing deep learning methods primarily focus on modeling image-level consistency through carefu...

Jingyi Xu, Xin Deng, Minglang Qiao et al. · 0 citations
Open access Sep 2026

Orthogonal Feature Decoupling and Hierarchical Cross-Scale Aggregation network for RGB-T semantic segmentation

Visible-thermal (RGB-T) imaging systems provide crucial complementary information for robust visual sensing and measurement in complex illumination conditions. However, effectively fusing multi-modal sensory data to achieve accurate semantic segmentation remains a key challenge. Most existing methods rely on heuristi...

Hong-Wei Liu, Yi-Sha Liu, Wei-Min Xue et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.