Jul 2026· International Conference Computing Methodologies and Communication· pp. 659-666· 0 citations· 19 references
Abstract
The fast advancement of deep neural networks has led to the escalation of hardware accelerator needs that achieve high functionality as they comply with strict requirements of power and latency, particularly in edge and embedded artificial intelligence. In this paper, the research introduce a low power, pipelined single-precision (32 bits) floating-point data path that is to be used in neural network accelerators compliant with the IEEE 754 single-precision standard. The suggested design uses multi-stage pipelining on addition, multiplication and accumulation units, which greatly decreases the critical path delays and enhances the overall throughput. Efficiency of power is also by ensuring that its techniques such as operand isolation, clock-conscious staging of pipelines and minimized switching activity in arithmetic units. The architecture has a combined optimization in latency, energy, and numerical accuracy, making it possible to infer the numerical accuracy of resource-constrained platforms in real-time. Simulations after synthesis show that the proposed data path has significant propagation delay and dynamic power improvements over the state-of-the-art non-pipelined floating-point implementations and can compute the accuracy needed by the deep learning workloads. Its scalable and modular design is flexible and can be easily adapted to other neural network designs. The findings demonstrate the strength of the targeted design towards addressing the increasing demand of high-performance, low-energy neural network hardware, which provides a viable approach to edge AI systems with severe demands on both power and performance.
The proposed Double MAC unit with dynamic precision scaling has showed twofold improvement in throughput and 15% improvement in power consumption, and proves advantageous in convolution layers, where greater precision is required for final classification and smaller precision in initial stage.
M. Jayasanthi, R. Kalaivani, K. P. Sampoornam et al.· Analog Integrated Circuits a...· 0 citations
The efficient implementation of the softmax is critical for optimizing transformer hardware accelerators. Unlike its role as a static, one-time classifier in convolutional neural networks (CNNs), softmax in transformers is core to achieving dynamic contextual awareness, generating attention weights that enable the mode...
Bangzheng He, Bang-Xin Qin, Han Wang et al.· IEEE Transactions on Very La...· 0 citations
The results indicate that exploiting the inherent parallelism of analog computation offers a promising pathway toward ultra-low-power AI inference, making the proposed architecture a potential alternative for energy-constrained edge applications.
Andrei Iliescu, O. N. Ionescu, Adrian Iosif· Electronics· 0 citations
This work proposes a hybrid SRAM-TCAM memory-centric deep neural network (DNN) architecture that integrates three tightly coupled subsystems and provides O(1) parallel associative inference.
Vandana Thakur, V. More, Abhishek Bhatt· Journal of King Saud Univers...· 0 citations
Multipliers dominate the critical path, power consumption, and silicon area of deep neural network (DNN) accelerators. This paper presents a high-efficiency 7-bit unsigned approximate multiplier tailored for DNN accelerators. Unlike conventional signed 8-bit designs—where the sign bit is handled separately via XOR—the...
H. Võ, T. Nguyen-Ly· IEEE International Conferenc...· 0 citations