Skip to content

TFNet: a triple-fusion network for multispectral pedestrian detection

Jul 2026 · Journal of Electronic Imaging (JEI) · Vol 35, pp. 043019 - 043019 · 0 citations · 44 references
Engineering

Abstract

Abstract. Occlusions and far distance pedestrians pose significant challenges for pedestrian detection, often leading to insufficient feature representation learned by models, which in turn results in degraded detection accuracy and a high miss rate. To address this issue, we propose a three-stage fusion multispectral pedestrian detection network named TFNet. The network employs a three-stage fusion strategy: first, the multiscale feature fusion attention module selectively enhances critical features within each modality, effectively focusing on discriminative regions of occluded and distant pedestrians. Subsequently, the transformer interactive fusion module establishes long-range dependencies and enables dynamic interaction across modalities, achieving deep semantic alignment and complementary information exchange. Finally, the pixel-adaptive feature fusion (PAFF) module performs pixel-level refinement and adaptive fusion of the interacted features, generating a more discriminative unified representation. Extensive experiments conducted on the public multispectral pedestrian detection KAIST dataset show that TFNet outperforms existing state-of-the-art algorithms. In particular, it achieves a steady and substantial reduction in miss rate on the occlusion subset and multiscale subset. Meanwhile, it obtains higher average precision on the FLIR and LLVIP datasets. All experimental results fully demonstrate that the proposed network can effectively alleviate missed detections of occluded and far-distance pedestrians, and is expected to present high application value in the field of pedestrian safety monitoring in public scenarios.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.