Jul 2026· International Conference on Machine Learning and Embedded Systems· Vol 14295, pp. 1429509 - 1429509-6· 0 citations· 11 references
Engineering
TL;DR
An end-to-end traffic scene recognition network based on the fusion of monocular camera images and corresponding road map top-down view data is proposed, with an overall recognition accuracy of 92.6%, outperforming the best single-input baseline Swin-Tiny by 3.3%.
Abstract
Accurate traffic scene recognition serves as a critical foundation for decision-making and safe driving in autonomous driving and intelligent transportation systems. Existing methods mostly rely on single visual data vulnerable to environmental variations, or vision-LiDAR fusion schemes with insufficient capacity to represent road topology and traffic semantic information, limiting recognition accuracy and robustness. To address these limitations, this paper proposes an end-to-end traffic scene recognition network based on the fusion of monocular camera images and corresponding road map top-down view data. We design a learnable cross-view spatial alignment module to eliminate perspective discrepancy, and a bidirectional cross-attention fusion module to enable deep bidirectional interaction between visual semantic and map topology features. Experiments on a self-built dataset covering five typical traffic scenes show that the proposed method achieves an overall recognition accuracy of 92.6%, outperforming the best single-input baseline Swin-Tiny by 3.3%. Ablation studies further validate the effectiveness of each core module.
With the large-scale implementation of scenarios such as logistics warehousing and park inspections, the demand for autonomous environmental perception by mobile robots continues to grow. Low-cost pure vision semantic perception has become a core technology for ensuring autonomous safe navigation of these robots. Tradi...
Jia-Wei Sun· Applied and Computational En...· 0 citations
A dual motion-model tracker that explicitly accounts for non-linear perspective transformations during vehicle approach is introduced, substantially improving temporal consistency over linear motion assumptions, and a semantic attribute classification pipeline that estimates occlusion level, readability, sign embeddedn...
Meda Lazar, S. Sridhar, Shashwata Gupta et al.· 0 citations
A unified multi-attribute framework based on a Vision Transformer, which is enhanced with lightweight, parameter-efficient adapters and exhibits consistent behavior–context relationships and demonstrates robustness under varied environmental conditions is proposed.
Surrounding scene awareness is a core component of self-driving techniques, and the 3D detection accuracy for occluded objects directly determines the system’s scenario adaptability and driving safety. To address the core problem of inter-object occlusions in traffic scenes that lead to reduced 3D detection accuracy an...
Jin Qi, Jian Wang· Journal of King Saud Univers...· 0 citations
Semantic segmentation for autonomous driving requires reliable detection of vulnerable road users (VRUs) despite heavy class imbalance. We introduce CLFTv2, a hierarchical camera-LiDAR fusion framework replacing global ViT attention with a Swin-based multi-scale encoder and a lightweight FPN-style residual decoder. Ope...
Toomas Tahves, Mauro Bellone, Raivo Sell· 0 citations
This work proposes a cascade optimization framework that systematically enhances feature representation and refines multimodal fusion, and introduces the Multi-Scale Contextual Fusion Module (MSCF) to reduce alignment bias.