Skip to content
Book Open access

Running AlphaFold3 on Distributed High-Throughput Computing Infrastructure: Scaling Workloads and Enabling Ultra-Large Predictions

Jul 2026 · Practice and Experience in Advanced Research Computing · pp. 1-5 · 0 citations · 9 references
Computer Science

TL;DR

dHTC is established as a viable—and in some regimes superior— execution model for data-intensive structural biology workflows and provide a general blueprint for deploying large, data-intensive applications on distributed cyberinfrastructure.

Abstract

AlphaFold 3 (AF3) enables atomic-resolution prediction of biomolecular complexes, driving rapidly growing demand across the life sciences. However, its ∼ 750,GB reference database has effectively confined production deployments to systems with shared parallel filesystems, creating a major barrier for scalability. Distributed high-throughput computing (dHTC) platforms offer vast, heterogeneous compute capacity, but fundamentally lack the shared data infrastructure assumed by AF3. We present a data-aware deployment of AF3 for dHTC, implemented on the Center for High Throughput Computing (CHTC) and the Open Science Pool (OSPool). The workflow is decomposed into a CPU-bound data pipeline that executes on nodes with locally staged, scheduler-advertised databases, and a GPU-bound inference pipeline that opportunistically scales across distributed resources. Using CUDA Unified Virtual Memory (UVM), we extend inference beyond physical GPU limits, enabling predictions of ultra-large complexes that exceed device vRAM. By elevating dataset locality to a schedulable resource via HTCondor ClassAds, we eliminate prohibitive per-job data transfers and enable efficient, federated execution. Beyond scaling throughput, we demonstrate that dHTC can support previously infeasible workloads. Together, these results establish dHTC as a viable—and in some regimes superior— execution model for data-intensive structural biology workflows and provide a general blueprint for deploying large, data-intensive applications on distributed cyberinfrastructure.

Read PDF

Similar papers

Preprint Aug 2026

Memory-, Circuit-, and Ansatz-Efficient VQLS for CFD on Hybrid Quantum-HPC Systems

Fluid dynamics workloads are dominated by repeated solves of large, structured linear systems, motivating the search for quantum acceleration. The Variational Quantum Linear Solver (VQLS) is a leading near-term candidate, but practical deployment on hybrid quantum--high--performance computing (HPC) systems faces three...

Chao Lu, M. G. Meena, Eduardo Antonio Coello Pérez et al. · 2 citations
#machine learning Preprint Sep 2026

COMPASS-ABS: Reducing Fragmentation in Shared GPU Clusters for Deep Learning Training Workloads

With the rapid advancement of deep learning technology, shared GPU clusters receive an increasing number of deep learning training (DLT) jobs. Yet resource fragmentation make such clusters underutilized and forces the DLT jobs running on them to endure long turnaround times. Extensive research has been devoted to quant...

Yu-Kai Zhou, Hong-Fan Wu · 0 citations
#machine learning Preprint Aug 2026

Deploying DeepSeek 175B Locally on a Single Consumer-Grade RTX 4060 Laptop with 32GB RAM for 200k-Scale Protein-Ligand Virtual Screening

This work validates the engineering feasibility of running industrial-scale trillion-parameter LLM-driven biomedical computing tasks on consumer hardware, establishing a new low-barrier paradigm for AI-powered early stage drug discovery.

Rui-Ya Xiao, Yi-Li Xu · 0 citations
Open access Sep 2026

A Scalable Distributed-Memory MPI Implementation of Smith-Waterman with Token-Passing Traceback

As genomic sequencing produces increasingly massive datasets, accurate local sequence alignment via the Smith-Waterman(SW) algorithm remains computationally prohibitive due to its space and quadratic time complexity O(mn). While parallelization addresses a path forward, existing MPI-based solutions present a critical b...

Maryam Fatima, Waqas Ali · 0 citations
Open access Sep 2026

LEO: Enabling Efficient Communication-Computation Pipeline for GNN via Hierarchical Caching

Training large-scale graphs with GNNs on multi-GPU platforms faces substantial feature loading overhead, leading to low resource utilization and inefficient training. Overcoming such communication bottlenecks is crucial, and current solutions achieve this by overlapping computation and communication through pipelining....

Jia-Qi Si, De-Zun Dong · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.