Skip to content
Conference

CIA: Composable Instruction-Based Accelerators

Sep 2026 · IEEE International Conference on Application-Specific Systems, Architectures, and Processors · pp. 41-48 · 0 citations · 26 references

Abstract

The use of custom instructions is a common approach to accelerate domain-specific problems. However, it is difficult to reuse instruction-based accelerators in new systems — for example, if multiple accelerators from past systems are merged into a single SoC, they may have colliding opcodes. Fixing this requires opcode reassignment (assuming this is possible), and all software needs to be recompiled. From this perspective, custom instruction-based accelerators are not portable and are therefore difficult to re-use. Another problem with such accelerators is providing OS-level access control and virtualization. While it is easy to make an accelerator exclusive to one process, it is challenging to share an accelerator in a multicore + multiprogrammed environment. Composable Instruction-based Accelerators, or CIA, solves these problems. First, it enables composability, where separately authored instruction-based accelerators can be merged into any future SoC and binary compatibility of software is maintained without recompilation. Second, it virtualizes the accelerator state, allowing multiple processes to share an accelerator while maintaining fully isolated state contexts under OS management. Through universal context save and restore that works with any accelerator, the OS provides access control and virtualized state management in multicore and multiprogrammed environments. The commercial impact is profound: CIA enables a marketplace of distinct Accelerator IP Vendors, who are domain experts, CPU Vendors, and SoC Vendors who create systems with arbitrary mixes of accelerators. CIA allows all hardware instances to execute the same software binaries, regardless of accelerator composition. This paper presents a case study, demonstrating how to deploy CIA into the RISC-V environment, yielding speedups through latency hiding, by switching between multiple accelerator contexts, and replication of accelerator units.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.