Skip to content
Conference

Robust traffic scene recognition via bidirectional cross-attention-based vision-map fusion

Jul 2026 · International Conference on Machine Learning and Embedded Systems · Vol 14295, pp. 1429509 - 1429509-6 · 0 citations · 11 references
Engineering

TL;DR

An end-to-end traffic scene recognition network based on the fusion of monocular camera images and corresponding road map top-down view data is proposed, with an overall recognition accuracy of 92.6%, outperforming the best single-input baseline Swin-Tiny by 3.3%.

Abstract

Accurate traffic scene recognition serves as a critical foundation for decision-making and safe driving in autonomous driving and intelligent transportation systems. Existing methods mostly rely on single visual data vulnerable to environmental variations, or vision-LiDAR fusion schemes with insufficient capacity to represent road topology and traffic semantic information, limiting recognition accuracy and robustness. To address these limitations, this paper proposes an end-to-end traffic scene recognition network based on the fusion of monocular camera images and corresponding road map top-down view data. We design a learnable cross-view spatial alignment module to eliminate perspective discrepancy, and a bidirectional cross-attention fusion module to enable deep bidirectional interaction between visual semantic and map topology features. Experiments on a self-built dataset covering five typical traffic scenes show that the proposed method achieves an overall recognition accuracy of 92.6%, outperforming the best single-input baseline Swin-Tiny by 3.3%. Ablation studies further validate the effectiveness of each core module.

View source

Similar papers

Open access Sep 2026

Deep Learning-Based Visual Semantic Perception Technology for Mobile Vehicles

With the large-scale implementation of scenarios such as logistics warehousing and park inspections, the demand for autonomous environmental perception by mobile robots continues to grow. Low-cost pure vision semantic perception has become a core technology for ensuring autonomous safe navigation of these robots. Tradi...

Jia-Wei Sun · 0 citations
Preprint Aug 2026

Multi-Modal Traffic Sign Detection with Semantic Attributes for Autonomous Driving

A dual motion-model tracker that explicitly accounts for non-linear perspective transformations during vehicle approach is introduced, substantially improving temporal consistency over linear motion assumptions, and a semantic attribute classification pipeline that estimates occlusion level, readability, sign embeddedn...

Meda Lazar, S. Sridhar, Shashwata Gupta et al. · 0 citations
Open access Jul 2026

Multi-Attribute Scene Context and Pedestrian Behaviour Recognition using PEFT-Tuned Vision Transformer Model for Autonomous Driving System

A unified multi-attribute framework based on a Vision Transformer, which is enhanced with lightweight, parameter-efficient adapters and exhibits consistent behavior–context relationships and demonstrates robustness under varied environmental conditions is proposed.

Tarun Reddi, Charvin Kusuma, Shamsad Parvin · 0 citations
Open access Aug 2026

MVXCC-NET: Cross-modal 3D detection of occluded objects based on dual-path information complementation and regional weight modeling

Surrounding scene awareness is a core component of self-driving techniques, and the 3D detection accuracy for occluded objects directly determines the system’s scenario adaptability and driving safety. To address the core problem of inter-object occlusions in traffic scenes that lead to reduced 3D detection accuracy an...

Jin Qi, Jian Wang · 0 citations
Preprint Sep 2026

CLFTv2: Efficient Camera-LiDAR Fusion for Semantic Segmentation via Hierarchical Feature Pyramids

Semantic segmentation for autonomous driving requires reliable detection of vulnerable road users (VRUs) despite heavy class imbalance. We introduce CLFTv2, a hierarchical camera-LiDAR fusion framework replacing global ViT attention with a Swin-based multi-scale encoder and a lightweight FPN-style residual decoder. Ope...

Toomas Tahves, Mauro Bellone, Raivo Sell · 0 citations
Aug 2026

Fadet: a fusion-aware 3D detection network with cascaded feature enhancement for small object detection in autonomous driving

This work proposes a cascade optimization framework that systematically enhances feature representation and refines multimodal fusion, and introduces the Multi-Scale Contextual Fusion Module (MSCF) to reduce alignment bias.

Chang-Hong Yu, Shaoshi Luo, Wen-Li Shen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.