Skip to content

Author

Jianyu Wang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

RoboPIM: A ReRAM-Based Accelerator for LLM-Based Robotics Applications via Dynamic Task Slicing

Large language models (LLMs) have demonstrated significant potential in improving robotics applications through natural language processing capabilities. Matrix multiplication dominates the computational workload in these LLM-based robotics systems, and resistive random access memory (ReRAM) provides notable performance benefits due to its inherent parallelism and energy efficiency. However, matrix multiplications in these applications pose dual heterogeneity challenges, encompassing both structural and scale heterogeneity. Structurally, these applications involve both standard dense matrix operations and block-structured sparse multiplications, where sparsity patterns vary significantly with robot topology. In terms of scale, matrix sizes range from just a few to thousands of dimensions, making large ReRAM crossbars designed for LLM inefficient for small-scale robotics computations. Existing acceleration approaches fail to address this dual heterogeneity, limiting their effectiveness for LLM-based robotics applications. In this work, we propose robotics applications-processing in memory (RoboPIM), the first low-energy ReRAM-based accelerator tailored for LLM-based robotics applications. robotics applications-processing in memory (RoboPIM) adapts to the diverse matrix multiplication requirements across robot platforms and LLM configurations. Our design features a ReRAM architecture with various crossbar sizes and implements a dynamic task slicing scheme that allocates matrix slices to appropriately sized crossbars. To maximize ReRAM’s parallel processing capabilities, we develop a two-stage scheduling strategy that efficiently manages matrix multiplications while enhancing hardware utilization. Comprehensive evaluations demonstrate that RoboPIM achieves $2.85\times $ and $6.85\times $ speedup across diverse robotic platforms and LLMs, respectively, outperforming state-of-the-art domain-specific solutions.

Wen-Jing Xiao, Jianyu Wang, Dan Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.