Large language models (LLMs) have demonstrated significant potential in improving robotics applications through natural language processing capabilities. Matrix multiplication dominates the computational workload in these LLM-based robotics systems, and resistive random access memory (ReRAM) provides notable performance benefits due to its inherent parallelism and energy efficiency. However, matrix multiplications in these applications pose dual heterogeneity challenges, encompassing both structural and scale heterogeneity. Structurally, these applications involve both standard dense matrix operations and block-structured sparse multiplications, where sparsity patterns vary significantly with robot topology. In terms of scale, matrix sizes range from just a few to thousands of dimensions, making large ReRAM crossbars designed for LLM inefficient for small-scale robotics computations. Existing acceleration approaches fail to address this dual heterogeneity, limiting their effectiveness for LLM-based robotics applications. In this work, we propose robotics applications-processing in memory (RoboPIM), the first low-energy ReRAM-based accelerator tailored for LLM-based robotics applications. robotics applications-processing in memory (RoboPIM) adapts to the diverse matrix multiplication requirements across robot platforms and LLM configurations. Our design features a ReRAM architecture with various crossbar sizes and implements a dynamic task slicing scheme that allocates matrix slices to appropriately sized crossbars. To maximize ReRAM’s parallel processing capabilities, we develop a two-stage scheduling strategy that efficiently manages matrix multiplications while enhancing hardware utilization. Comprehensive evaluations demonstrate that RoboPIM achieves $2.85\times $ and $6.85\times $ speedup across diverse robotic platforms and LLMs, respectively, outperforming state-of-the-art domain-specific solutions.
Wen-Jing Xiao, Jianyu Wang, Dan Chen et al.· IEEE Transactions on Compute...· 0 citations
If AI is to support human cognitive growth, design must move beyond answer provision and efficiency maximization toward the organization of productive human-AI relations: relations that challenge users’ initial assumptions while providing support appropriate to the task and the user's level of expertise.
Xiaokun Wu, Min Chen, Giancarlo Fortino· Big Data and Cognitive Compu...· 0 citations
Fusing digital twins (DTs) with world models (WMs) promises a leap from reactive monitoring to proactive control. However, deploying high-fidelity WMs faces a fundamental cognitive gap: the immense computational demand of foundation models conflicts with the resource constraints of 6G edge nodes, while privacy regulations hinder the centralization of raw sensory data required for training. To bridge this gap, this paper presents HongAvg, a hierarchical on-demand cognitive split federated learning framework designed as cognitive infrastructure for next-generation DTs. By establishing a national-basin-edge three-tier architecture, HongAvg introduces foundation models to resource-constrained edges via split computing, offloading heavy cognitive reasoning while preserving data privacy. We propose a dual-stream semantic consistency mechanism to align edge interactions with the foundation model's cognition, ensuring that distributed cognitive primitives serve as valid inputs for global state estimation. Validated on a heterogeneous benchmark, HongAvg serves as a prototype for industrial cognitive computing, improving accuracy in visual monitoring by up to 8.9% and reducing edge memory usage by approximately 75%. This work provides the scalable architectural prerequisite necessary for evolving static DTs into proactive WM-driven multi-agent orchestrated intelligent systems.
Yue Wang, Jixuan Xie, Yusheng Lin et al.· IEEE Transactions on Network...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.