Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 7566-7576· 0 citations· 13 references
TL;DR
RecCompl is a comprehensive model compilation system that enables efficient model compilation of industrial scale DLRMs with PT2 and introduces a configuration-based user interface that decouples compilation settings from model code, allowing fine-grained control without intrusive changes.
Abstract
Deep Learning Recommendation Models (DLRMs) play a key role to power real-world recommendation and ranking, yet their growing complexity has made production deployment increasingly challenging. While PyTorch 2 (PT2) offers promising performance and productivity gains through automated model compilation, its initial release lacked critical features needed for DLRM adoption. In this work, we present RecCompl, a comprehensive model compilation system that enables efficient model compilation of industrial scale DLRMs with PT2. RecCompl addresses key compatibility issues by extending operator coverage, minimizing graph breaks, avoiding unnecessary recompilation, and generalizing graph transformation. Besides, to meet the high requirement on model exploration, we introduce a configuration-based user interface that decouples compilation settings from model code, allowing fine-grained control without intrusive changes. Despite these improvements, efficiency gaps remain in achieving a production-ready compilation system. To close them, we introduce systematic designs and engineering optimizations that enhance compilation time, memory management, and online deployment reliability. RecCompl is widely adopted, and delivers up to 60% higher training throughput while achieving substantially lower compilation latency, consistent performance across varying memory budgets, and stable online deployment.
Scaling model capacity has emerged as an effective approach to overcoming performance bottlenecks in industrial recommender systems. However, repeatedly training larger dense models from scratch demands substantial data and time, while their growing computation conflicts with the strict serving budgets of industrial sy...
Rui-Hao Zhang, Bo Chen, Xiao Wang et al.· 0 citations
The engineering details of ADEPT are presented, a framework that integrates representative adequacy techniques, including neuron-coverage-based metrics, surprise adequacy, input distribution coverage, boundary coverage, and source- and model-level mutation score, under a consistent execution workflow.
Yidi Kao, Shawn Burnham, Tommi Rose Fahy et al.· 0 citations
The design of transformer-based Large Language Models (LLMs) is being radically changed through new architectures that are able to overcome scalability limitations of previous designs, including Mixture-of-Experts (MoE), Multi-Head Latent Attention (MLA), and Multi-Token Prediction (MTP). As an open-weighted model rele...
Yassine Zouhdi, B. Hdioud· EPJ Web of Conferences· 0 citations
The embedding tables in Deep Learning Recommendation Models (DLRMs) require significant memory capacity and bandwidth but relatively lower computing power, making it economically inefficient to scale by adding more GPUs solely to meet memory requirements. Recent advances in Compute Express Link (CXL) and near-data proc...
Zheng Wang, Zhong-Kai Yu, Kai-Jian Wang et al.· ACM Transactions on Architec...· 0 citations
PALRec is proposed, a parameter-preserving augmentation framework that equips an LLM with recommendation capabilities while keeping its original parameters fixed and consistently outperforms fully fine-tuned counterparts in recommendation accuracy while preserving the LLM’s pre-trained knowledge.
Hyunsoo Na, Minseok Gang, Sang-goo Lee et al.· ACM Transactions on Informat...· 0 citations
Large language models (LLMs) have demonstrated strong capabilities in recommendation tasks such as item, sequence, conversational recommendation, and explanation generation. However, LLM weights are typically shared across all users. Adapting these models to individual users remains a fundamental challenge that require...
Kanishka Dandeniya, C. Dasanayaka, Daswin de Silva et al.· Proceedings of the 20th ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.