Open access
Jul 2026
Exploring High-Bandwidth Flash for Modern LLM Inference: Opportunities and Challenges
Analysis of HBF-based LLM-serving systems under diverse system configurations and operating scenarios shows that HBF can significantly improve the batch size, throughput, and flexibility of LLM-serving systems while reducing the minimum GPU requirements, but realizing these benefits critically depends on sustaining HBM-comparable read bandwidth and requires significant endurance improvements.
D. Son, Yonggon Park, Hyunuk Cho et al.
· IEEE computer architecture l... · 2 citations