Enhanced mask-region-based segmentation with parallel CNN for object detection in traffic scenarios
Abstract
Millions of lives are lost annually in traffic accidents. To reduce these risks, it is vital to assist vehicles in identifying important objects that may pose a threat. Past research has focused on individual participants, often ignoring their interrelations, thus making it less effective in complex traffic scenarios. To address these issues, this study proposes a novel normalized parallel convolutional neural network for object detection in traffic scenarios (NPC-ODT). ODT starts with an input video, which is broken down into individual frames for in-depth processing. Each frame is first subjected to preprocessing using the Wiener filter—a technique that minimizes noise while maintaining essential image features, thereby improving overall image clarity. After preprocessing, segmentation is carried out using a modified bottleneck attention-based mask recurrent convolutional neural network (MBA-MRCNN), which efficiently separates key regions of interest from the surrounding environment. Once segmentation is complete, a hierarchical feature extraction process is applied. This involves analyzing the updated center pixel in hierarchy of skeleton (UC-HoS) to understand their structural form, extracting color features to distinguish between different object types, and using motion estimation to capture dynamic behavior across frames. The collected features are then input into a NPC neural network (NPCNN), which performs the task of identifying and classifying the objects present in each frame. The NPC-ODT scheme attained the highest accuracy rate at 0.937, precision at 0.907, and MCC at 0.859, significantly outperforming established methods. The final result is a set of accurately detected and labeled objects ready for further analyses, such as tracking, decision-making, or behavioral prediction.