RecDM: Efficient Training System for Large-Scale Recommendation Models on Disaggregated Memory
The embedding tables in Deep Learning Recommendation Models (DLRMs) require significant memory capacity and bandwidth but relatively lower computing power, making it economically inefficient to scale by adding more GPUs solely to meet memory requirements. Recent advances in Compute Express Link (CXL) and near-data proc...