Implementation of a MobileNet-Based Accelerator for Real-Time Image Segmentation on FPGA
Abstract
Real-time image segmentation is essential for edge-based intelligent systems. Deep learning models are efficient for image segmentation compared to conventional techniques. Among other deep neural networks (DNN), MobileNet convolution neural network (CNN) model utilizes depthwise separable convolutions to reduce computational complexity. This work developed a field programmable gate array (FPGA) implementation of MobileNet model using hardware-aware optimizations such as fixed-point quantization and batch normalization fusion. The model is implemented using Vitis high-level synthesis (HLS) and validated on FPGA to demonstrate efficient hardware acceleration. Experimental results demonstrate a pixel accuracy of 93.67%, dice score of 0.4638, and intersection over union (IoU) of 0.3211. In this work, two architectures are developed, one is hardware efficient and other one is fast. The hardware implementation achieves low latency and efficient resource utilization, making it suitable for edge deployment. The proposed system demonstrates an effective trade-off between segmentation accuracy and hardware efficiency for real-time applications.