Skip to content
Book Open access

Algorithm--Hardware Co-Design of Spiking Transformers for In-Memory Neuromorphic Edge Vision

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 35 references

Abstract

Reliable, ultra-low-power neuromorphic vision on edge devices requires attention mechanisms that combine accuracy with parallel, memory-efficient execution. Current Spiking Transformers suit neuromorphic data but retain the \(\mathcal {O}(N^2D)\) complexity of standard self-attention and weak temporal modeling, limiting their efficiency on in-memory edge hardware. We propose the Temporal Hadamard Transformer (THT), a spiking model co-designed with hardware for in-memory edge vision. Its Temporal Hadamard Attention (THA) combines a lightweight Temporal Processing Unit (TPU) for short-term temporal fusion with a binary Mask&Add operator. By replacing Query–Key–Value computation with accumulate-and-fire logic that maps efficiently onto memristor crossbars, THA reduces attention complexity to \(\mathcal {O}(ND)\) and enables highly parallel in-memory execution. On CIFAR10-DVS, THT achieves a peak Accuracy-to-Energy ratio of 157.06. THT consumes 0.51–2.19 mJ per inference. Its lowest-energy configuration uses up to 16 × lower energy than the compared models, while the largest configuration reaches 81.8% accuracy. We further design a detailed hardware deployment scheme and evaluate its robustness through non-ideal memristor crossbar simulations. Accounting for conductance drift and IR drop, the simulated deployment maintains 81.1% accuracy. Thus, THT integrates sparse algorithms with parallel hardware execution, providing a practical path toward real-time edge systems under realistic resource constraints and physical device non-idealities.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.