FPGA-SoC Implementation of 3D-HEVC Encoder Based on MD-CNN
Abstract
Convolutional Neural Networks (CNNs) are the most common deep learning architecture used for video pro-cessing enhancement. Particularly, the Multi-Deep Convolutional Neural Network (MD-CNN) model, embedded into the 3D-HEVC encoder, was able to extract the optimal CTU partition structure efficiently in the depth map and successfully substitute the time-consuming rate-distortion optimisation (RDO) full traversal search, reducing the coding complexity heavily. Nevertheless, MD-CNN requires significant memory resources and is highly computationally intensive. Field-Programmable Gate Arrays (FPGAs), particularly the emerging FPGA–SoC technology, are recognised as the most promising platforms for accelerating CNNs due to their high performance, energy efficiency, and reconfigurable nature. In this study, complementary technologies have been combined and deployed to allow high-performance 3D-HEVC video processing. By exploiting the inherent parallelism of FPGA architectures, MD-CNN processing throughput was significantly improved, execution latency was reduced, and energy efficiency was enhanced, while maintaining the coding performance and compression efficiency of the 3D-HEVC video encoder. The experimental results demonstrate that the proposed co-design achieves an on-chip power consumption of 1.891 W.