This work presents a hardware–software co-design whose 16-bit tag uses odd parity and fail-safe class encoding, which combines a variable-precision extent, an aligned CRC checker, and an exact-bounds micro-cache.
Abstract
Graphics processing units (GPUs) underpin high-performance computing, but device partitioning does not ensure that an in-context pointer remains within its allocation. We present a hardware–software co-design whose 16-bit tag uses odd parity and fail-safe class encoding. It combines a variable-precision extent, an aligned CRC checker, and an exact-bounds micro-cache. For 200,000 log-uniform requests from 16 B to 1 TiB, mean Class 4 fragmentation is 1.638%, versus 27.733% for power-of-two encoding. The seven-bit aligned signature has full GF(2) rank and a 1/128 non-adaptive collision rate; its deterministic window is exactly one through three regions and tight at four. Exhaustive testing rejects every one-bit tag corruption, while two-bit analysis demonstrates why parity is not adversarial authentication. A SAT-equivalent endpoint rewrite reduces Class 3 generic depth from 60 to 24 levels. A fail-closed 32-lane Class 4 topology detects nonuniform active-lane tags in hardware; for uniform tags, it reduces generic CMOS cost from 155,200 to 72,872 transistor equivalents (53.05%) and depth from 67 to 61 levels. Official SASS traces provide an analytical exposure bound rather than native simulation; a separate pre-layout 45 nm mapping is reported only as a timing sensitivity experiment.
In the modern world, 128-bit binary arithmetic is essential for achieving high numerical accuracy but remains expensive to implement in processors. This research presents an optimized, reduced-clock-cycle approach for performing 128-bit floating-point calculations on Field-Programmable Gate Arrays (FPGAs) using the SRM...
N. P, Thirumalaiswamy V., H. S. et al.· International Journal of Com...· 0 citations
Object detection at the edge requires a difficult balance among detection accuracy, deterministic latency, memory bandwidth, and energy consumption. Existing binarized accelerators replace multipliers with XNOR and population-count logic, but many designs use a fixed binary datapath or select precision only at the laye...
Budidha Sriman and D. Sateesh· International Journal of Adv...· 0 citations
Large language model (LLM) inference transfers model weights and activations for every generated token, making memory traffic and its energy cost part of the decode path. BitNet b1.58 represents its low-bit weights by ternary values and uses integer activations. However, this arithmetic does not match conventional int8...
Across three decoder architectures on an H100 the authors measure that parse, not copy, holds 64-72% of device-resident decode time; that bounding back-reference chain depth - provable, and costing 0.006% in ratio - moves latency by at most 2.8% and, for the file's own latency spike, provably by nothing at all.
Modern GPU kernels increasingly stress the instruction supply path, while fixed instruction containers can leave substantial footprint slack. This paper presents an automated encoding-synthesis framework that treats instruction layout as a constrained slot-assignment problem over a validated instruction-form field spec...
Cryptographic hash functions over integers modulo a prime play a decisive role in the efficiency and security of proof systems for computational integrity. Early designs focused on compact arithmetic circuits and efficient software execution, primarily targeting general-purpose CPUs rather than hardware accelerators. T...
Luca Campa, Thomas De Cnudde, Al Kindi et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.
Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…
AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.