Skip to content

Is INT8 Portable? A Cross-Platform Measurement Study of Quantized Inference on Embedded and Automotive Accelerators

Sep 2026 · 0 citations · 43 references
Computer Science

Abstract

Eight-bit integer (INT8) post-training quantization is the default recipe for edge deployment, under a widely held assumption: INT8 makes inference faster at a small, predictable accuracy cost, and a model quantized once can be carried to any target. We test that assumption with a controlled measurement study across seven hardware classes -- ARM and x86 CPUs, a discrete GPU, an NVIDIA Jetson AGX Orin iGPU and its NVDLA cores, and two vendor NPUs (Qualcomm Hexagon HTP, DEEPX DX-M1) -- holding the ONNX artifact and the quantization scales fixed so the integer kernel or ISA is the only free variable. Portability fails on three axes. (1) The sign of the INT8 speedup is set by the CPU's dot-product ISA (ARM dotprod/SDOT, x86 VNNI): cores that have it speed up by up to 2.1x, cores that lack it slow down by 1.7x, for the identical model and runtime. (2) INT8 outputs are not portable, and the rule is an invariance rather than a gradient: FP32 predictions are bit-identical for every pair (1000/1000), while INT8 predictions agree 1000/1000 exactly when two targets share an integer kernel and 958-965/1000 whenever they do not -- independent of whether the boundary is CPU<->CPU or CPU<->accelerator, and invisible to top-1 accuracy, which is preserved. (3) Vendor NPUs own quantization: a bring-your-own QDQ graph fails silently on one NPU (external scales ignored, accuracy 0.75 ->0.005 while it compiles, profiles and runs without error) and loudly on the other (the compiler refuses the graph), so only the vendor's native path yields a correct engine. We further show that edge-NPU latency regimes are set by output/device-to-host transfer size rather than compute, and locate the transition with a fixed-compute sweep. We release the scripts and 32 reports."Quantize once, deploy anywhere"is unsafe for embedded and automotive deployment, where per-input determinism and redundancy matter.

View source

Similar papers

#machine learning Preprint Aug 2026

The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inference

Two GPU kernels implementing the same scaled INT8 GEMM interface are usually treated as interchangeable but cross-implementation FP8 GEMM shows a different signature: both the prevalence and the magnitude of differences grow with reduction depth, while the INT8 fraction stays at parts per million and within one spacing...

Teng-Ruei Chen · 2 citations
Conference Sep 2026

Bridging the DSP Gap: A Streamed, Multiplier-Less Edge AI Operator via Logarithmic Co-Design

Low-cost FPGA SoCs can be limited by DSP availability even when logic remains available. We present an edge-AI operator that co-designs selected-layer logarithmic quantization with a synthesizable shift-and-add datapath and AXI-compatible integration. Across classification, segmentation, and detection, short Log-QAT re...

Chang-Yan Liu, Ming Xiao, Sha-Sha Feng · 0 citations
Preprint Sep 2026

Quantifying the Effect of HCLs on a Fixed-Microarchitecture MXFP4 Accelerator

This paper compares the most widely used HCLs using the same fixed design, the OCP MXFP4 block dot product, a quantization primitive at the heart of edge Physical-AI inference, implemented as a single 12-stage, II=1 pipeline.

D. Passaretti, Sajjad Tamimi, Nicola Dall'Ora · 0 citations
#artificial intelligence Preprint Sep 2026

NPU Accelerator: Quantized Real-Time Vehicle Detection on PYNQ-Z1 Using FINN

This paper presents the design, optimization, implementation, and on-board validation of a neural processing unit (NPU) accelerator for real-time vehicle detection on the resource-constrained Xilinx Zynq XC7Z020 device of the PYNQ-Z1 board. The work follows a hardware/software co-design methodology that combines quanti...

Daniel G. Gutiérrez, Antonio Cuesta, Jorge D. Fe et al. · 0 citations
#natural language process... Preprint Sep 2026

I-Parakeet: Integer-Only Conformer ASR on Mobile NPU

In this paper, we propose I-Parakeet, an integer-only implementation of NVIDIA's Parakeet-CTC (0.6B parameters) that runs on a smartphone NPU without any floating-point operator or CPU fallback. Modern Conformer ASR models are hard to deploy on edge devices because of their size, and quantized models still fall back to...

Taichi Nishimura · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.