The embedding tables in Deep Learning Recommendation Models (DLRMs) require significant memory capacity and bandwidth but relatively lower computing power, making it economically inefficient to scale by adding more GPUs solely to meet memory requirements. Recent advances in Compute Express Link (CXL) and near-data proc...
Zheng Wang, Zhong-Kai Yu, Kai-Jian Wang et al.· ACM Transactions on Architec...· 0 citations
How to encode and evaluate spatial intelligence in foundation models remains an open challenge. Existing approaches often rely on textual proxies and VQA-style evaluation to assess visual-spatial intelligence (VSI), which can obscure geometric structure, encourage linguistic shortcuts, and hinder attribution to genuine...
Guan-Lin Wu, Bo-Yan Su, Yang Zhao et al.· IEEE Transactions on Pattern...· 0 citations
Multi-LoRA serving is how one base model becomes thousands of specialized variants, one adapter per user, task, or agent, and the deployments can hold 1000-plus adapters. Serving them is hard because the workload inverts what GPUs provide: terabytes of memory against only tens of TFLOPS, and because every published sys...
Zhong-Kai Yu, O. Venkatachalam, Zheng Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.