Traffic Sign Image Segmentation using U-Net Architecture Based on Convolutional Neural Network
Abstract
Accurate pixel-level segmentation of traffic signs in natural scenes is a critical precursor to reliable sign recognition in intelligent transportation systems. This study investigates the effectiveness of the U-Net architecture for semantic segmentation of traffic signs using a dataset comprising 1,750 training subset and 300 testing subset. The model was trained for 20 epochs with the Adam optimizer (learning rate 0.0002, batch size 2), exhibiting stable convergence: training loss decreased from 0.137729 to 0.05989, while the dice coefficient improved from 0.899644 to 0.956839, and binary IoU rose from 0.894176 to 0.94945. Evaluation on a held-out test set of 300 images demonstrated strong generalization, yielding a dice coefficient of 0.930935, binary IoU of 0.907808, precision of 0.90499, recall of 0.981679, and accuracy of 0.953572. The recall–precision asymmetry indicates a mild tendency toward over-segmentation. The consistent covariation between the dice coefficient and binary IoU across both training and testing phases affirms the internal reliability of the evaluation metrics employed in this study. Qualitative analysis revealed high-fidelity segmentation of polygonal signs, with the model faithfully reproducing the sharp vertices and overall silhouette of both the octagonal and the rhomboid sign. In contrast, the circular sign exhibited a localized boundary failure, characterized by conspicuous concave flattening along its upper-left contour that caused the predicted rim to appear clipped relative to the smooth ground-truth circumference. These findings establish U-Net as a robust segmentation foundation for downstream traffic sign recognition, while motivating future work on architectural refinements to improve boundary-level precision for curvilinear signs and the adoption of recall loss functions to further regulate the recall–precision trade-off and mitigate over-segmentation tendencies.