Transformer attention mechanisms pose significant scalability challenges due to quadratic complexity in sequence length, and existing accelerators remain bottlenecked by dense arithmetic and data movement. This paper proposes CAMformer, a hardware accelerator that reinterprets attention as an associative memory operati...
Tergel Molom-Ochir, Benjamin F. Morris, Mark Horton et al.· IEEE Transactions on Circuit...· 0 citations
Edge intelligence promises responsive, private, and energy-efficient sensing without continual dependence on remote compute. This demands convolutional neural network (CNN) accelerators that deliver substantially higher throughput and energy efficiency than conventional digital pipelines while preserving high accuracy....
Mark Horton, Changwoo Park, Tergel Molom-Ochir et al.· International Symposium on L...· 0 citations
Extreme compression techniques, including vector quantization (VQ) and input-dependent sparsity, can significantly reduce the memory footprint of large language models (LLMs). However, a key challenge remains in translating such compression into practical efficiency. On conventional systolic-array-based accelerators, V...
Haoxuan Shan, Cong Guo, Bo-Wen Duan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.