Aug 2026· Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication· 0 citations· 49 references
Computer Science
TL;DR
CAISA is introduced, a composable AI systems architecture that enables disaggregated memory expansion for large-scale AI workloads using CXL-based shared memory and coupling memory isolation with software-managed data orchestration decouples compute and memory resources while preserving efficient data movement, providing a scalable and cost-efficient foundation for future AI infrastructure.
Abstract
Modern state-of-the-art AI systems are increasingly built as monolithic supernodes integrating large numbers of specialized accelerators with proprietary high-bandwidth interconnects. These systems provision compute, memory, and networking resources in fixed ratios at design time. As AI workloads evolve, their resource demands increasingly diverge from these static configurations, leading to underutilization, limited scalability, and high operational cost. Composable systems based on disaggregated resources offer a more flexible alternative by allowing memory and compute capacity to be scaled independently without replicating an entire supernode. We introduce CAISA, a composable AI systems architecture that enables disaggregated memory expansion for large-scale AI workloads using CXL-based shared memory. CAISA goes beyond conventional memory disaggregation by jointly designing hardware support, runtime mechanisms, and workload mapping to provide contention-free shared-memory access across CPUs, accelerators, and CXL memory devices. Its key mechanism is workload-aware memory isolation, which maps shared memory through reserved address spaces to control data visibility, avoid hardware coherence overheads, and enable fine-grained pipelined data movement between compute and memory devices. By coupling memory isolation with software-managed data orchestration, CAISA decouples compute and memory resources while preserving efficient data movement, providing a scalable and cost-efficient foundation for future AI infrastructure.
This work proposes SAI, a mechanism that virtualizes shared memory into the L2 cache to improve GPU performance for AI applications and introduces an L2 cache management strategy that integrates associativity-based virtual page allocation and a replacement information table, reducing page-swapping overhead while preser...
Hanqing Li, Tie-Jun Li, Sheng Ma et al.· ACM Transactions on Design A...· 0 citations
This paper identifies effective I/O-computation overlap as a key requirement for fully exploiting the AI data center stack, and outlines future research directions for next-generation analytical database architectures.
Ji-Gao Luo, Nils Boeschen, Muhammad El-Hindi et al.· Datenbank-Spektrum· 0 citations
Memory disaggregation provides key-value stores larger memory capacity at low cost. Emerging compute express link (CXL) enables efficient memory disaggregation. It, however, dramatically slows down the system performance as disaggregated memory accesses are considerably slower than local memory accesses. This paper pre...
Chencheng Ye, Yuanchao Xu, Xipeng Shen et al.· ACM Transactions on Architec...· 0 citations
The growing disparity between processor core scaling and memory bandwidth has exposed the physical and economic limits of processor-centric database architectures. While Compute Express Link (CXL) and other emerging technologies enable a necessary shift toward memory-centric, disaggregated topologies, it also introdu...
Yi Jiang, Hamish Nicholson, Anastasia Ailamaki· Datenbank-Spektrum· 0 citations
The evolution of supercomputer architecture has undergone several transformative phases
since the 1960s, yet existing surveys have not adequately captured the engineering trade-offs
that defined each generation. This paper presents a critical survey of supercomputer
architecture from early vector machines to contemp...
Oraye Godspower· International Journal of Com...· 0 citations
Evaluation using cycle-accurate hardware-software co-simulation demonstrates that CHIPSMORE effectively sustains high throughput across varying model sizes, context lengths, and batch sizes while maintaining favorable power scaling.