Skip to content
Book Open access

RecCompl: Efficient Model Compilation for Industrial Scale Recommendation Models with PyTorch 2

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 7566-7576 · 0 citations · 13 references

TL;DR

RecCompl is a comprehensive model compilation system that enables efficient model compilation of industrial scale DLRMs with PT2 and introduces a configuration-based user interface that decouples compilation settings from model code, allowing fine-grained control without intrusive changes.

Abstract

Deep Learning Recommendation Models (DLRMs) play a key role to power real-world recommendation and ranking, yet their growing complexity has made production deployment increasingly challenging. While PyTorch 2 (PT2) offers promising performance and productivity gains through automated model compilation, its initial release lacked critical features needed for DLRM adoption. In this work, we present RecCompl, a comprehensive model compilation system that enables efficient model compilation of industrial scale DLRMs with PT2. RecCompl addresses key compatibility issues by extending operator coverage, minimizing graph breaks, avoiding unnecessary recompilation, and generalizing graph transformation. Besides, to meet the high requirement on model exploration, we introduce a configuration-based user interface that decouples compilation settings from model code, allowing fine-grained control without intrusive changes. Despite these improvements, efficiency gaps remain in achieving a production-ready compilation system. To close them, we introduce systematic designs and engineering optimizations that enhance compilation time, memory management, and online deployment reliability. RecCompl is widely adopted, and delivers up to 60% higher training throughput while achieving substantially lower compilation latency, consistent performance across varying memory budgets, and stable online deployment.

Read PDF

Similar papers

Preprint Sep 2026

Inherit4Rec: Parameter Inheritance for Efficient Scaling of Recommendation Models

Scaling model capacity has emerged as an effective approach to overcoming performance bottlenecks in industrial recommender systems. However, repeatedly training larger dense models from scratch demands substantial data and time, while their growing computation conflicts with the strict serving budgets of industrial sy...

Rui-Hao Zhang, Bo Chen, Xiao Wang et al. · 0 citations
Book Open access Aug 2026

ADEPT: A Unified Framework for Deep Learning Test Adequacy

The engineering details of ADEPT are presented, a framework that integrates representative adequacy techniques, including neuron-coverage-based metrics, surprise adequacy, input distribution coverage, boundary coverage, and source- and model-level mutation score, under a consistent execution workflow.

Yidi Kao, Shawn Burnham, Tommi Rose Fahy et al. · 0 citations
Conference Open access 2026

DeepSeek-V3: Architecture and Optimizations-A Practical Review

The design of transformer-based Large Language Models (LLMs) is being radically changed through new architectures that are able to overcome scalability limitations of previous designs, including Mixture-of-Experts (MoE), Multi-Head Latent Attention (MLA), and Multi-Token Prediction (MTP). As an open-weighted model rele...

Yassine Zouhdi, B. Hdioud · 0 citations
Open access Sep 2026

RecDM: Efficient Training System for Large-Scale Recommendation Models on Disaggregated Memory

The embedding tables in Deep Learning Recommendation Models (DLRMs) require significant memory capacity and bandwidth but relatively lower computing power, making it economically inefficient to scale by adding more GPUs solely to meet memory requirements. Recent advances in Compute Express Link (CXL) and near-data proc...

Zheng Wang, Zhong-Kai Yu, Kai-Jian Wang et al. · 0 citations
#large language models Review Open access Sep 2026

PALRec: Large Language Model-Based Sequential Recommendation With Parameter-Preserving Augmentation

PALRec is proposed, a parameter-preserving augmentation framework that equips an LLM with recommendation capabilities while keeping its original parameters fixed and consistently outperforms fully fine-tuned counterparts in recommendation accuracy while preserving the LLM’s pre-trained knowledge.

Hyunsoo Na, Minseok Gang, Sang-goo Lee et al. · 0 citations
Book Open access Sep 2026

GradSup: Gradient Superposition for Personalised and Scalable LLM Recommendation

Large language models (LLMs) have demonstrated strong capabilities in recommendation tasks such as item, sequence, conversational recommendation, and explanation generation. However, LLM weights are typically shared across all users. Adapting these models to individual users remains a fundamental challenge that require...

Kanishka Dandeniya, C. Dasanayaka, Daswin de Silva et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.