An Ensemble Recognition Model Based on Cross-Modal Fusion of HRRP and Infrared Image
Abstract
Infrared (IR) image data is effective under low-light and cluttered conditions, offering stable imaging and highlighting thermal features, though it lacks detailed spatial resolution. High-resolution range profile (HRRP) data, on the other hand, captures target structural signatures with strong penetration and all-weather capability but suffers from low interpretability due to its 1-D form. To address the limitations and complementarity of these modalities, this article proposes a multimodal fusion framework that combines IR and HRRP data through a unified recognition strategy. The approach includes data-level preprocessing, feature-level modeling with handcrafted and deep representations, and decision-level fusion via a linear support vector machine (SVM) ensemble. MobileNetV2, a lightweight multilayer perceptron (MLP), and a 1-D convolutional neural network (1D-CNN) are employed for modality-specific classification, and their outputs are integrated to enhance recognition robustness. The experimental results on a simulated paired dataset comprising 23 target subcategories grouped into four major recognition classes demonstrate that the proposed model outperforms single-modality baselines, particularly in complex and noncooperative scenarios, confirming its effectiveness and potential for practical applications. Repeated evaluations further distinguish recognition of held-out views from generalization to unseen target subcategories.