Aug 2026· International Symposium on Low Power Electronics and Design· pp. 1-7· 0 citations· 39 references
Computer Science
TL;DR
This work proposes a flattening methodology that preserves GHRR's matrix-based encoding while executing training and inference directly in vector space, functionally equivalent to FHRR inference yet free of permutation logic.
Abstract
Symbolic reasoning at scale requires representations that are both expressive and hardware-efficient. Hyperdimensional Computing (HDC) provides a promising brain-inspired framework, encoding concepts as high-dimensional vectors composed through efficient algebraic operations. However, the widely used Fourier Holographic Reduced Representation (FHRR) relies on commutative binding, which cannot encode order or directionality. Prior FPGA/ASIC accelerators recover this through costly permutation logic, incurring irregular memory access and high control overhead. Generalized Holographic Reduced Representations (GHRR) instead redefine binding as non-commutative matrix multiplication, naturally encoding sequences, graphs, and hierarchies without permutations—but at the cost of turning every element into a matrix, introducing compute and storage overhead that limits scalability. In this work, we propose a flattening methodology that preserves GHRR's matrix-based encoding while executing training and inference directly in vector space, functionally equivalent to FHRR inference yet free of permutation logic. We further design a custom 28 nm ASIC that fuses binding and similarity into a unified complex-valued datapath, with dual-DMA streaming and runtime normalization for accurate inference. Against a PyTorch baseline on an NVIDIA RTX 4090 GPU, the prototype delivers 1.36×-1.56× higher throughput and 16.2×-18.6× better energy efficiency.
Matrix multiplications dominate the inference cost of modern transformer-based vision models, yet existing efficiency techniques such as post-training quantization and mixed-precision inference are largely limited to the small set of fixed-width formats (INT4, INT8, BF16, and FP16) supported by conventional accelerator...
Hao-Ran Jin, Kang-Qi Zhang, Ji-Rong Yang et al.· 0 citations
Computer vision requires intense tensor operations, primarily matrix multiplications, imposing substantial computation demands. Using hardware such as GPU and ASICs for acceleration offers a viable solution. Their application at edge, however, can be constrained by complexity in the computing architecture and incompati...
Jingfang Pei, Lekai Song, Songwei Liu et al.· Advances in Materials· 0 citations
BaKron is an efficient solver that combines anti-diagonal parallelism with a recursive divide-and-conquer construction that matches the cubic scaling of GPTQ while exploiting richer curvature information.
Autonomous systems that learn and explore over long horizons face a problem. Standard methods scale poorly in the number of observations, n, precluding sustained operation on bounded hardware. We show that compositional, high-dimensional vector representations inspired by neural computation address these constraints. W...
P. M. Furlong, Nicole Sandra-Yaffa Dumont, Rika Antonova et al.· Nature Communications· 0 citations
Fast matrix multiplication algorithms keep the product fixed and search for a cheaper way to evaluate it. We instead ask whether a Transformer's learned projections can use a different, cheaper product altogether. Building on an associative-algebra construction that replaces ordinary matrix multiplication with a sparse...
The reformulation preserves the benefits of WTConv while substantially reducing its execution time and memory footprint, removing the systems overhead that previously limited its practical efficiency.
Amit Aflalo, Shahaf E. Finder, Roy Amoyal et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.