Skip to content
Preprint

APEX: A Dual-Sparsity Accelerator for Precise and Efficient SNN Inference

Aug 2026 · 0 citations · 15 references
Computer Science

TL;DR

APEX is presented, a dual-sparsity SNN inference accelerator that integrates the PASC-IF neuron into the LoAS hardware framework, and guarantees mathematical equivalence between the converted SNN and the source ANN, thereby achieving ANN-equivalent accuracy at significantly reduced timesteps.

Abstract

Spiking Neural Networks (SNNs) have emerged as an energy-efficient alternative to Artificial Neural Networks (ANNs), leveraging sparse accumulate operations in the place of power-hungry multiply-and-accumulate operations. ANN-SNN conversion is a widely adopted approach to realize deep SNNs with accuracy comparable to that of ANNs. The Quantization-Clip-Floor-Shift (QCFS) activation minimizes conversion error, yet requires a large number of inference timesteps to match the source ANN accuracy on real-world vision datasets. PASCAL addresses this by proposing the Precise ANN-SNN Conversion Integrate-and-Fire (PASC-IF) neuron, which guarantees mathematical equivalence between the converted SNN and the source ANN, thereby achieving ANN-equivalent accuracy at significantly reduced timesteps. Despite this algorithmic advancement, the hardware implications of deploying the PASC-IF neuron remain unexplored. In this work, we present APEX, a dual-sparsity SNN inference accelerator that integrates the PASC-IF neuron into the LoAS hardware framework. The three-stage PASC-IF datapath is realized as a fully combinational circuit with no additional latency cost. APEX exploits dual sparsity in both input spikes and weights through a fully temporal-parallel dataflow, enabling efficient sparse computation and reduced memory traffic. Across all evaluated models, the PASC-IF neuron on average achieves up to 3% higher accuracy than the standard IF neuron, with a power overhead of only 1.3%-5.4%, an area overhead of 2.1%-2.7%, and 40% energy reduction for best accuracy configurations.

View source

Similar papers

Preprint Nov 2025

NeuroFlex: Lossless Element-Level ANN-SNN Co-Execution for Efficient Sparse Inference

Sparse DNN accelerators specialize in ANN or SNN execution, leaving energy or latency on the table when workload characteristics vary within a layer. Hybrid accelerator designs that switch modes at layer or tile granularity suffer from low PE utilization since one core type idles whenever the other is active. NeuroFlex...

Varun Manjunath, P. Ramesh, Gopalakrishnan Srinivasan · 0 citations
Preprint Aug 2026

Reducing ANN-SNN Conversion Error via Residual Membrane Potential Alignment

This work analyzes flaws of conventional conversion pipelines from residual membrane potential statistics and proposes a novel conversion strategy combining dynamic initial potential tuning and feature enhancement, which generalizes to ReLU CNNs, ANN Transformers, and multi-threshold SNN variants.

Zirui Chen, Zihan Huang, Tong Bu et al. · 0 citations
Book Open access Aug 2026

A Differentiable Simulator for Optimizing Time-Domain Analog CNN Accelerators

Edge intelligence promises responsive, private, and energy-efficient sensing without continual dependence on remote compute. This demands convolutional neural network (CNN) accelerators that deliver substantially higher throughput and energy efficiency than conventional digital pipelines while preserving high accuracy....

Mark Horton, Changwoo Park, Tergel Molom-Ochir et al. · 0 citations
Open access Aug 2026

SAD-SNN: Spatial-Activation Distillation for High-Performance Spiking Neural Networks

A Spiking Neural Network (SNN) is a kind of brain-inspired and event-driven network, which is becoming a promising energy-efficient alternative to Artificial Neural Networks (ANNs). In recent years, SNN methods have been successfully applied in the fields of electromagnetic signal processing and image signal processing...

Chongxiao Qu, Qian Zhang, Chenxiao Dou et al. · 0 citations
#machine learning Preprint Aug 2026

Event-Driven Language Models with Sparse Neural Activity for Neuromorphic Hardware

This work introduces a method that induces sparse neural activity in heavily quantized linear-attention models with minimal performance loss, and positions sparse, quantized linear-attention models as a natural fit for deploying LLMs on event-driven multi-core platforms.

Simon Richter, Ruhai Lin, Jason Yik et al. · 0 citations
Jul 2026

The Sparsity Ceiling: Where Spiking Networks Can and Cannot Trade Activity for Energy

It is argued the energy dividend of sparsity is not a property of SNNs but of the task, and the ceiling is formalized with an information-theoretic bound and confirmed: the floor rises with memory load, falls with state width, and (refuting a naive memory-only reading) rises with task difficulty.

Zeyu Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.