Skip to content
Open access

TFE-Fusion: Tri-Modality Feature Enhanced Fusion for Robust Object Detection

2026 · IEEE Open Journal of Vehicular Technology · Vol 7, pp. 2593-2605 · 0 citations · 43 references

Abstract

Robust and reliable object detection under adverse conditions remains a critical challenge for automated driving systems (ADS). The performance of RGB-based (visible-spectrum) cameras degrades in poor lighting conditions, whereas the performance of RGB-Long-Wave Infrared (LWIR) fusion architectures remains limited in adverse weather and low thermal contrast scenarios due to partial spectral coverage. At the same time, Short-Wave Infrared (SWIR) has the potential to address some of these gaps but remains under-utilised in ADS. To address these limitations, we propose a novel tri-modality fusion architecture, TFE-Fusion, that simultaneously leverages information from RGB, SWIR, and LWIR modalities. Our architecture employs a feature-level fusion strategy that incorporates pixel-level weighting and a spectral attention mechanism, enabling dynamic fusion of complementary features from all three modalities. A YOLOv8-based detection head, modified for multi-modal input streams, is used for efficient and robust inference. Extensive experiments on the Multispectral Object Detection (MOD) dataset demonstrate that the proposed method outperforms the RGB-LWIR baseline by 2.48 percentage points in mAP@0.5:0.95. The most significant performance gains are observed for the vehicle class, where SWIR features compensate for low thermal contrast in LWIR imagery by capturing distinct material reflectance properties. These results validate the effectiveness of the proposed TFE-Fusion architecture for enhancing detection performance, particularly under low-light and low-thermal-contrast conditions.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.