Jul 2026· Journal of Intelligent Computing and Networking· Vol 2, pp. 1-13· 0 citations· 49 references
TL;DR
The proposed HDAM (Hierarchical Domain Adaptation with Multi-Level Attention), an enhanced domain-adaptive detection framework built upon Faster R-CNN, effectively bridges the modality gap and achieves improved cross-domain detection performance compared to recent DAOD methods.
Abstract
Cross-modal domain adaptation for RGB-to-Thermal object detection (DAOD) remains a challenging task due to the large modality gap, feature distribution mismatch, and limited supervision on the thermal domain. Conventional domain adaptation detectors often fail to capture consistent semantic representations across modalities, resulting in unstable adversarial learning and suboptimal detection accuracy. To address these challenges, we propose HDAM (Hierarchical Domain Adaptation with Multi-Level Attention), an enhanced domain-adaptive detection framework built upon Faster R-CNN. HDAM progressively aligns cross-domain features from low- to high-level representations through a hierarchical attention mechanism. Specifically, it comprises three key components: (1) Entropy-Guided Cross-level Attention (ECA), which leverages discriminator-derived entropy responses to reweight pixel- and mid-level features for more stable cross-modal encoding; (2) Multi-branch Global Semantics and Discriminator (MGSD), which performs global adversarial alignment and integrates multi-scale spatial information through attention-enhanced branches, while introducing an auxiliary image-level semantic consistency loss; and (3) Instance-Context Alignment (ICA), which fuses ROI-level and contextual embeddings and applies dual instance-level adversarial losses for precise fine-grained adaptation. Extensive experiments on benchmark RGB-Thermal datasets demonstrate that HDAM effectively bridges the modality gap and achieves improved cross-domain detection performance compared to recent DAOD methods.
A Multi-modal Interaction Enhanced Segment Anything Model (MIE-SAM) that reconfigures SAM's image encoder into a weight-sharing dual-branch image encoders, and translates the fused features into the fine-grained saliency map in an entirely prompt-free, end-to-end manner.
Ze Li, Ying-Ying Zhang, Shuai Zhang et al.· Neural Networks· 0 citations
Visual object tracking (VOT) in real-world scenarios necessitates a high degree of adaptability to overcome the distributions discrepancies inherent in multi-domain environments. While standard deep learning models excel in well-conditioned laboratory datasets, their performance typically degrades when encountering adv...
Volodymyr Husiev, I. Vergunova· Automation, Control, and Inf...· 0 citations
A Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) is proposed to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives.
Hu Lin, Zhi-Wei Fu, Xiu-Mei Chen et al.· Remote Sensing· 0 citations
Remote sensing object detection faces three challenges: extreme scale variation, arbitrary rotational orientations, and complex intermingled backgrounds. Although fusion of Convolutional Neural Networks (CNNs) and Transformers combines spatial precision with global modeling, it faces two limitations: (i) feature distri...
Mo Zhou, Yue Zhou, Kai Song· PLoS ONE· 0 citations
Robust and reliable object detection under adverse conditions remains a critical challenge for automated driving systems (ADS). The performance of RGB-based (visible-spectrum) cameras degrades in poor lighting conditions, whereas the performance of RGB-Long-Wave Infrared (LWIR) fusion architectures remains limited in a...
Muhammad Usama, Imad Ali Shah, Roshan George et al.· IEEE Open Journal of Vehicul...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.