This work proposes a high-performance NAND flash-based ISC architecture with enhanced activation buffering, and introduces a distributed dataflow approach for the NAND-PIM array that maximizes computational parallelism by employing efficient intra-plane data mapping.
Abstract
In-storage computing (ISC) is considered a next-generation memory architecture for its ability to relieve the data bottleneck between the host and the memory. While the required resources of large language models (LLMs) have increased significantly in recent years, the memory density has not scaled accordingly. Recently, several works have studied NAND flash-based processing-in-memory (NAND-PIM) schemes to exploit the high density of the memory. However, they do not address the dataflow/buffer for the intermediate values, so a simple method is to deal with the values in the slow flash memory array. To overcome such a limitation, we propose a high-performance NAND flash-based ISC architecture with enhanced activation buffering. Instead of using the very slow flash memory array for the intermediate values, our architecture buffers the values in a fast DRAM subsystem. This approach effectively handles the high-latency penalties when activations are programmed into slower TLC NAND flash. We also introduce a distributed dataflow approach for the NAND-PIM array. This approach maximizes computational parallelism by employing efficient intra-plane data mapping. The results show that our proposed architecture achieves significant performance improvements, reducing the inference latency by up to 85% compared to the baseline.
Advances in flash memory technology have increased the number of bits stored per cell, significantly reducing the cost of solid-state drives (SSDs). As a result, SSDs using high-density flash have become common, though they suffer from lower performance and endurance than low-density flash. Modern SSDs exploit the abil...
Cha-Nu Yu, Jongseok Kim, Euiseong Seo· IEEE Access· 0 citations
This paper presents a novel memory controller architecture and a RISC-V instruction set extension to optimize MLC NVM write operations by balancing speed and retention time, and introduces a fast-store instruction in RISC-V to increasing write performance while addressing retention limitations.
Mina Ibrahim, M. Shokry, Lokesh Siddhu et al.· Journal of Circuits, Systems...· 0 citations
Conventional computing architectures are reaching their scalability limits, while their energy demands increase rapidly. A bottleneck is the separation of memory and processing units, which requires a continuous data transfer. This memory wall increases the power consumption and limits the processing speed at the same...
L. Brackmann, Tobias Ziegler, N. Kopperberg et al.· npj Unconventional Computing· 0 citations
FLINT is proposed, a workload-driven HBF substrate for capacity-scalable LLM inference that integrates HBF as a memory-capacity tier alongside HBM while addressing three adoption challenges.
Geraldo F. Oliveira, Arash Tavakkol, Xiang-Yu Zhu et al.· 1 citation
Hybrid flash storage combines large-capacity highdensity flash memory with high-performance low-density flash memory, providing excellent cost-effectiveness. Existing data placement strategies for hybrid flash storage typically employ hotness-based data migration relying on a twotier architecture. This approach not onl...
Han Yan, Dingcui Yu, Yanyun Wang et al.· IEEE Non-Volatile Memory Sys...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.