Feature-Level Multimodal Fusion for Integrated Vehicle Condition and Performance Prediction
Abstract
The accurate prediction of vehicle condition and performance is crucial for improving the safety of road transport systems, self-propelled vehicles, and sustainable intelligent transport systems. To address the challenge, a multimodal feature-level fusion and Hierarchical Multimodal Transformer Network (HMT-Net) are proposed to achieve integrated vehicle condition and performance prediction. The proposed framework leverages the nuScenes v1.0 dataset, which includes synchronized camera images, LiDAR point clouds, radar measurements, and keyframes annotated with objects and other safety-critical information from real-world driving scenarios. Heterogeneous sensor data is first preprocessed using modality-specific normalization to reduce noise and normalize the data. A Swin Transformer is used to extract discriminative visual features from camera images, a Voxel Transformer (VoTr) is used to learn 3D geometric representations from LiDAR point clouds, and a Bidirectional Gated Recurrent Unit (Bi-GRU) is used to capture temporal motion information from radar signals. A novel Adaptive Cross-Modal Attention Fusion (ACMAF) module is proposed to integrate the extracted features to generate a comprehensive and informative feature representation. A Transformer Encoder is then used to model the interactions between fused features followed by a Multilayer Perceptron (MLP) for vehicle condition and performance prediction.