Skip to content
Open access

An Integrated Framework for High-Precision and Low-Power Computing

Jul 2026 · African Journal Of Applied Research · Vol 12, pp. 597-627 · 0 citations · 24 references

TL;DR

A unique combination of Posit arithmetic, iterative division, SIMD parallelisation, and pipelining optimisation into a single framework that addresses the shortcomings of IEEE-754 floating-point arithmetic and can be used effectively in high-performance computing systems, embedded processors, and AI accelerators.

Abstract

Purpose: The goal of this research is to create a unified computational framework that addresses the shortcomings of IEEE-754 floating-point arithmetic by enhancing the equilibrium among precision, dynamic range, and hardware resource efficiency. Design / Methodology / Approach: The proposed method integrates Posit number representation, Newton-Raphson-based division, SIMD-based parallel execution, and pipeline optimisation techniques into a single hybrid computational framework. Posit arithmetic is used to increase accuracy with fewer bits, and Newton-Raphson iterations are used instead of hardware dividers to reduce latency. The data were analysed via comparative performance evaluation through simulation experiments. Performance ratios, percentage improvements, and graphs were used to analyse the results and assess improvements in accuracy, speed, parallel processing, and energy efficiency. Research Limitation: The present implementation and assessment are confined to simulated or regulated computational settings and particular nonlinear workloads. The framework has not been validated across all floating-point workloads or various real-time hardware platforms, which may impact its generalisability. Findings: The input/design variables, processing variables, and performance/output variables cut the bit-width by 20–30% compared to IEEE-754 and make numbers more accurate by up to 35%. Precision up to 10−610^{-6}10−6 is reached within three Newton–Raphson iterations, which cuts division latency by 40–50%. SIMD execution increases throughput by 60–70% and energy efficiency by about 20%. Optimising the pipeline cuts latency by 30% to 40%. In general, the design reduces hardware space by up to 20% and improves energy efficiency by 18%. Practical Implication: The suggested framework can be used effectively in high-performance computing systems, embedded processors, and AI accelerators to achieve higher numerical accuracy, use less hardware space and power, and increase data throughput without complicating the architecture. Social Implication: The proposed arithmetic framework supports sustainable computing practices by enabling more accurate, energy-efficient computing. Originality / Value: This study introduces a unique combination of Posit arithmetic, iterative division, SIMD parallelisation, and pipelining optimisation into a single framework. The overall benefits to precision, latency, hardware area, and energy efficiency make it a unique and scalable solution for next-generation computing.

Read PDF

Similar papers

Open access Aug 2026

An Open-Source Evaluation Framework for RISC-V Co-Design-Based Decimal Arithmetic

Hardware–software co-design is a balanced strategy for computationally intensive algorithms such as decimal computing. It can provide several Pareto points for the development of embedded systems in terms of hardware cost and performance. In this study, we propose an efficient and accurate evaluation framework for deci...

R. Mian, M. Inoue · 0 citations
Open access Aug 2026

Design and Implementation of a High-Performance RISC-V CPU IP Core for Multiple Cryptographic Algorithms

A hierarchical multiplier architecture consisting of partial-product generation, 4–2 compressor-tree reduction, and a look-ahead adder is developed to improve multiplication throughput and reduce latency compared with conventional Wallace-tree implementations.

Y.-J. Cheng, Cheng-Rui Yin, K.-X. Tang et al. · 0 citations
Preprint Aug 2026

DiffPower: GPU-Accelerated Differentiable Switching Power Analysis and Optimization

DiffPower translates design netlists into a PDK-agnostic bytecode representation, enabling analytical gradient computation via reverse-mode automatic differentiation, achieving up to a speedup over single-threaded CPU propagation on the largest evaluated design, with the GPU advantage growing with design scale.

Isaac Jacobson, Zhengjie Zhao, Rashmi Mehrotra et al. · 0 citations
Conference Aug 2026

An FPGA-Based Unified Processing Element for INT8/Binary Quantization and its Scalable Array Design

The widespread deployment of deep neural networks on edge devices faces a severe imbalance between computational demand and available power, while devices frequently switch between low-power standby and highperformance detection modes. Existing general-purpose processors, graphics processing units, and fixed-precision...

Zi-Qing Mai, Zhan-Peng Jiang · 0 citations
Open access Sep 2026

Optimization of Parallel Number Theoretic Transform Algorithms for Multi-Core Digital Signal Processors

The Number Theoretic Transform (NTT), a finite-field variant of FFT, is critical in cryptography and digital signal processing, but its efficiency on FT-M7032 Digital Signal Processors(DSPs) remains suboptimal due to memory bottlenecks and architectural constraints. This paper proposes a tailored optimization framework...

Ding-Xing Xie, Xiao-Chuan Hu, Lin Peng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.