Skip to content

ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts

Jul 2026 · arXiv.org · Vol abs/2607.17074 · 0 citations · 36 references
Computer Science

TL;DR

ThAME is proposed, a three-dimensional (3D) heterogeneous multi-chiplet architecture for MoE inference that employs Ferroelectric Field-Effect Transistor-based non-volatile and DRAM-based volatile memory chiplets with a co-designed compute mapping strategy that aligns the distinct computational profiles of attention mechanisms and expert routing.

Abstract

Mixture of Experts (MoE) architectures have emerged as a dominant paradigm for scaling Large Language Models (LLMs). However, MoE inference on conventional hardware is constrained by three fundamental bottlenecks. These encompass the massive memory bandwidth required to fetch non-contiguous expert weights, the non-deterministic scatter-gather traffic generated by input-dependent token routing, and the tail-latency dependency imposed by synchronous expert output aggregation. To address these challenges, we propose ThAME, a three-dimensional (3D) heterogeneous multi-chiplet architecture for MoE inference. ThAME employs Ferroelectric Field-Effect Transistor (FeFET)-based non-volatile and DRAM-based volatile memory chiplets with a co-designed compute mapping strategy that aligns the distinct computational profiles of attention mechanisms and expert routing. Furthermore, we design a specialized Network-on-Chip communication backbone optimized to mitigate the bottlenecks associated with non-deterministic token routing traffic across the combinatorial space of input-dependent MoE traffic patterns. Experimental results demonstrate that ThAME outperforms state-of-the-art counterparts by up to 15.7x in terms of speedup and improves energy efficiency by up to 9.8x.

View source

Similar papers

Preprint Sep 2026

LLM Inference on IMC-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism

LLM inference has become an essential service, yet it imposes unprecedented demands on memory bandwidth, computational density, and communication efficiency. While IMC is a promising solution to the memory wall issue, the heterogeneous data dynamicity of LLM requires complementary resources to handle intermediate data...

Yimin Wang, Yue Jiet Chong, Xuan-Yao Fong · 0 citations
Open access Sep 2026

HDA-MoE: Hybrid Parallelism and Dynamic, Adaptive Scheduling for Mixture-of-Experts with 3D Near-Memory Processing

HDA-MoE is presented, a framework that optimizes MoE execution on NMP architectures through hybrid parallel deployment and runtime scheduling and integrates an offline hybrid parallel mapping algorithm with an online dynamic and adaptive scheduling mechanism to reduce communication overhead while improving computation...

Hao-Chen Huang, Shu-Zhang Zhong, Sheng-Xuan Qiu et al. · 0 citations
Book Open access Aug 2026

HyNA: Taming Tail Latency in MoE Training with Hybrid Switch Silicon

HyNA is a fully serverless aggregation system that eliminates dedicated parameter-server nodes by leveraging a novel hardware-software co-designed switch architecture, and reduces synchronization time by up to 1.6X compared to dynamic INA baselines, without compromising bit-level model accuracy.

Yang Liu, Tianxiang Liu, Hai-Peng Yao · 0 citations
Review Aug 2026

AI Hardware Accelerators for Large Language Models: Architectures and the Memory Wall

This review surveys the full landscape of AI hardware accelerators for LLMs, including general-purpose GPUs, custom ASICs, reconfigurable FPGAs, processing-in-memory and near-memory architectures, and emerging neuromorphic and photonic approaches across cloud and edge deployment.

Siddharth Patel, Rohit Singh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.