Optimization of Parallel Number Theoretic Transform Algorithms for Multi-Core Digital Signal Processors
Abstract
The Number Theoretic Transform (NTT), a finite-field variant of FFT, is critical in cryptography and digital signal processing, but its efficiency on FT-M7032 Digital Signal Processors(DSPs) remains suboptimal due to memory bottlenecks and architectural constraints. This paper proposes a tailored optimization framework: a four-step strategy decomposes large 1D NTT into 2D transforms, outperforming traditional six-step methods by reducing memory access overhead; a VLIW/SIMD-aware microkernel with loop unrolling and register reuse boosts computation; vectorized Montgomery modular multiplication supports 16 parallel streams to address the lack of division instructions; and double buffering with pipelining overlaps computation and data transfer. Experiments show a 90.6× speedup over baselines for large-scale NTTs, with ablation studies validating each optimization. This work offers insights for NTT acceleration on DSPs and data-intensive algorithm optimization on heterogeneous processors.