A Balancing Optimization And Automation In Hardware Design Considering an Application To Pipelined Booth Multiplier
Abstract
Many High-Level Synthesis (HLS) tools have been developed recently, targeting faster design cycles and improved productivity. These tools employ different methodologies to structure the data flow and deal with pipelining, loop unrolling, and other architectural alternatives, which can drastically impact design performance. Hence, there is a significant need for a comparative performance analysis among the RTL generated by these tools to guide designers toward the most suitable choice based on specific optimization criteria and the implementation technology. In this context, this manuscript provides a comparative analysis of the Hardware Description Language (HDL) code generated from MATLAB Simulink, Python, and C source files using Simulink HDL Coder, MyHDL, and AMD Vitis, respectively. This analysis involves computing the trade-off between design productivity, considering time and effort, and hardware performance, considering required resources/area, speed, and power consumption, for both FPGA and ASIC implementations. To accomplish this investigation, a phase shifter using pipelined Booth multipliers is selected as the design under test with discussing the digital beamforming as one of its applications. This design is suitable because its performance is sensitive to resource constraints, and its architecture is not trivial but has suitable complexity and flexibility. The design is represented in four versions: a direct Verilog HDL code, MATLAB Simulink, Python code, and C code. The results of the direct Verilog representation are taken as the reference for all comparisons. The RTL of these versions are simulated for functionality verification using the Vivado design suite then synthesized and implemented on the Virtex UltraScale+ VCU118 evaluation platform, xcu9p-flga2104-2L-e, for FPGA implementation. On the other hand, for measuring the suitability of the different HDL coders for ASIC realization, ASIC synthesis of these RTL versions was performed utilizing the TSMC 65 nm standard cell library. The comparisons reveal that MyHDL offers a good balance of flexibility (through parameterization and pipelining), and performance. This tool achieves the lowest resource utilization (61.14% reduction in LUTs number and 42.88% reduction in registers number compared to the Direct Verilog), and also 75% less power consumption for FPGA realization. However, the delay increases slightly by 6.8%. Similarly, for the ASIC implementation, the MyHDL-generated RTL shows superior performance with the most area-efficient design (88.3% less than the direct Verilog), the lowest power consumption (91.49% lower than the direct Verilog. But with lower maximum clock frequency (3.66% less than the direct Verilog). Conversely, the AMD Vitis-generated RTL exhibits the worst performance metrics in the FPGA and the ASIC implementation, considering the area and max clock frequency with enhancement in the consumed power by 46.9% for ASIC implementatio. For the MATLB Simulink, it shows the min delay in case of FPGA implementation, which reaches 27.85% reduction than the direct Verilog.