Skip to content

RACE-AIMC: Selective Inference for Heterogeneous Analog In-Memory Accelerators at the Edge

Sep 2026 · 0 citations · 14 references
Computer Science

TL;DR

RACE-AIMC (Risk-Aware Certified Ensemble for AIMC), a framework that resolves the choice to trust a single chip blindly with statistics rather than guesswork and matches the accuracy of a clean digital baseline while cutting modeled energy use.

Abstract

Analog in-memory computing (AIMC) speeds up neural-network inference by doing the arithmetic directly inside a memory array, instead of shuttling weights back and forth between memory and a processor. This saves energy, but the physical devices that store the weights are imperfect: programming errors, electrical noise, limited-resolution converters, and outright broken cells all distort the computation, and every physical chip is distorted in its own way. A designer with several such chips available faces an uncomfortable choice: run all of them and combine the answers (safe, but wasteful of energy), or trust a single chip blindly (cheap, but with no guarantee on how often it is wrong). This paper introduces RACE-AIMC (Risk-Aware Certified Ensemble for AIMC), a framework that resolves this choice with statistics rather than guesswork. Offline, RACE-AIMC studies a pool of physical accelerators, picks the single best one for a given energy budget, and computes a mathematically exact upper bound on how often that accelerator will be wrong when it chooses to answer. Online, only that one accelerator is switched on; a lightweight check decides whether to accept its answer or defer to a fallback. In our simulations using a noisy weight mapping and multiple independent test runs, every certified bound stayed under a 10% error target (mean bound 7.83% +- 0.89%, with 70.88% +- 0.98% of inputs answered directly). The resulting system matches the accuracy of a clean digital baseline while cutting modeled energy use by 69.02% relative to always running every accelerator in the pool.

View source

Similar papers

Preprint Sep 2026

The World Model Hardware Accelerator

Diffusion transformers invert the arithmetic that autoregressive decoding made familiar. There is no token-by-token recurrence: every denoising step is a full-sequence forward pass over static shapes, so the entire schedule is known at compile time and the only serial dimension is the step count itself. We exploit that...

Shashank Chaurasia · 0 citations
Open access Sep 2026

Versat-AI: An ONNX-to-SoC Compiler for Model-Agnostic CGRA Edge Inference

Edge inference on resource-constrained embedded nodes demands accelerators that are energy-efficient and compact. This paper presents Versat-AI, an open-source compiler that accepts a standard Open Neural Network Exchange (ONNX) model and generates a complete, synthesisable RISC-V System-on-Chip (SoC) with an embedded...

R. Teixeira, J. Rodrigues, Jaime Aguiar et al. · 0 citations
Preprint Sep 2026

Quantifying the Effect of HCLs on a Fixed-Microarchitecture MXFP4 Accelerator

This paper compares the most widely used HCLs using the same fixed design, the OCP MXFP4 block dot product, a quantization primitive at the heart of edge Physical-AI inference, implemented as a single 12-stage, II=1 pipeline.

D. Passaretti, Sajjad Tamimi, Nicola Dall'Ora · 0 citations
#edge computing Preprint Sep 2026

FlexSpIM: An Event-Based Digital Compute-In-Memory Accelerator with Flexible Operand Resolution and Layer-Wise Hybrid Stationarity

FlexSpIM, a digital CIM architecture supporting arbitrary operand resolution and shape within a unified storage for weights and neuron states, is introduced, enabling a layer-level hybrid weight- and output-stationary dataflow, maximizing operand reuse and reducing costly on- and off-chip data movement during SNN execu...

Nicolas Chauvaux, Adrian Kneip, Charlotte Frenkel · 0 citations
#machine learning Preprint Sep 2026

Is INT8 Portable? A Cross-Platform Measurement Study of Quantized Inference on Embedded and Automotive Accelerators

Eight-bit integer (INT8) post-training quantization is the default recipe for edge deployment, under a widely held assumption: INT8 makes inference faster at a small, predictable accuracy cost, and a model quantized once can be carried to any target. We test that assumption with a controlled measurement study across se...

Yu-Len T. Shin · 0 citations
Conference Sep 2026

QXL - A Scalable Hardware Framework for Tabular Q-Learning Inference on RISC-V

Tabular Q-learning is a foundational reinforcement learning (RL) algorithm used in embedded decision-making, yet its inference phase suffers from latency-sensitive memory accesses, creating severe bottlenecks in real-time control systems. To address this, we present QXL, a hardware-software codesigned inference acceler...

C. R, Pruthvi Parate, P. K. et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Microsoft Research Blog Aug 20, 2026

Broadening access to Skala creates a faster path to predictive DFT 

Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. The post Broadening access to Skala creates a faster path to predictive DFT  appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.