Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators

The rapid growth in machine learning workloads has fueled the proliferation of custom accelerator architectures. Designed from the ground up, these accelerators often expose programming models that are distinct from GPUs. While hyperscalers and AI chip startups continue to innovate in this space, achieving broad operator coverage to support diverse models remains a major challenge. Additionally, an easy-to-use, high-level kernel programming language is important for rapid iteration of models and kernels. Triton, together with TorchInductor, addresses these issues on GPUs, but its viability on accelerators with different programming models has yet to be established. In this work, we present the first production-scale application of Triton on a custom ML accelerator, MTIA-2i, developed by Meta. To support MTIA-2i, we develop a new compiler backend that targets it, introduce enhancements to TorchInductor code generation, and propose minimal language extensions that expose MTIA-specific architectural features. We demonstrate that Triton-MTIA kernels achieve performance competitive with expert-tuned C++ implementations. Leveraging these development efficiency gains, we successfully deployed manually written and Inductor-generated Triton kernels in production across approximately 60 different model types, accounting for 50% of layers and 47% of non-GEMM execution time for these models. Our results provide compelling evidence that DSLs like Triton can bridge the programming model gaps between ML frameworks, kernels, and custom accelerators, enabling rapid innovation and efficient deployment at scale.

Haishan Zhu, Domi Yan, Michael Levesque-Dion et al. · 0 citations
Open access Aug 2026

A Multi-stage Precision Stratification (MPS) Framework for Navigating Adjuvant Immunotherapy in Hepatocellular Carcinoma After Resection

Background: Recurrence rates following curative resection for hepatocellular carcinoma (HCC) remain persistently high, benefit from adjuvant immunotherapy varies substantially across patients, and the field currently lacks a standardized framework to characterize the postoperative host immune contexture. Purpose: To propose and validate a Multi-stage Precision Stratification (MPS) framework and evaluate its value in prognostic stratification and prediction of immunotherapy response. Methods: The Immune Health Index (IHI = S + R - E) integrating immune surveillance (S), immune exhaustion (E), and immune reserve (R) was constructed to define four immune phenotypes. Prognostic value was assessed in four public HCC cohorts (n=931) with single-cell transcriptomic validation (GSE140228, 61,690 cells); a blood-count-based clinical version cIHI_v8 was constructed in the Qinghai QPHCC cohort (n=490 survival analysis). Results: IHI was an independent protective prognostic factor in TCGA-LIHC (multivariate HR=0.795, P=0.034); four-cohort random-effects meta-analysis yielded HR=0.818 (95% CI: 0.696-0.961), I-squared=31.4%. QPHCC cIHI_v8 multivariate HR=0.452, HR=0.715 after ALBI adjustment; Bayesian evidence synthesis yielded BF_10=1280 for cIHI_v8 (>100 constitutes Decisive evidence), whereas the 4-cohort meta BF_10=2.19 (Anecdotal). Following NLP-based reverse stage derivation (n=490, achieving full AJCC/BCLC stage coverage from 0%), IHI remained significant after AJCC adjustment (HR=0.8642, P=0.000079), IHI provided positive incremental C-index across all stage-adjusted models; stratified analysis showed the strongest effect in early-stage (AJCC I-II: HR=0.8109, P<0.0001) and MVI-negative patients (HR=0.8538, P=0.0020). Bootstrap 1000x resampling: median HR=0.8646 (95% CI: 0.7985-0.9443), all iterations yielded HR<1. Conclusions: The MPS framework provides a mechanism-driven biological stratification tool for adjuvant immunotherapy in post-resection HCC, moving from "fixed-protocol extrapolation" to "immune contexture navigation."

Z. Dang, J. Dan, W. Su et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.