Algorithm--Hardware Co-Design of Spiking Transformers for In-Memory Neuromorphic Edge Vision
Abstract
Reliable, ultra-low-power neuromorphic vision on edge devices requires attention mechanisms that combine accuracy with parallel, memory-efficient execution. Current Spiking Transformers suit neuromorphic data but retain the \(\mathcal {O}(N^2D)\) complexity of standard self-attention and weak temporal modeling, limiting their efficiency on in-memory edge hardware. We propose the Temporal Hadamard Transformer (THT), a spiking model co-designed with hardware for in-memory edge vision. Its Temporal Hadamard Attention (THA) combines a lightweight Temporal Processing Unit (TPU) for short-term temporal fusion with a binary Mask&Add operator. By replacing Query–Key–Value computation with accumulate-and-fire logic that maps efficiently onto memristor crossbars, THA reduces attention complexity to \(\mathcal {O}(ND)\) and enables highly parallel in-memory execution. On CIFAR10-DVS, THT achieves a peak Accuracy-to-Energy ratio of 157.06. THT consumes 0.51–2.19 mJ per inference. Its lowest-energy configuration uses up to 16 × lower energy than the compared models, while the largest configuration reaches 81.8% accuracy. We further design a detailed hardware deployment scheme and evaluate its robustness through non-ideal memristor crossbar simulations. Accounting for conductance drift and IR drop, the simulated deployment maintains 81.1% accuracy. Thus, THT integrates sparse algorithms with parallel hardware execution, providing a practical path toward real-time edge systems under realistic resource constraints and physical device non-idealities.