Deep Learning for Proteins: a series of 10 interactive notebook modules that introduce fundamental machine-learning concepts, guide users through training machine-learning models for protein-related tasks, and ultimately present cutting-edge protein structure prediction and design pipelines are developed.
Abstract
Computational methods for predicting and designing biomolecular structures are increasingly powerful. Although previous approaches relied on physics-based modeling, modern tools (e.g., AlphaFold2 in CASP14) leverage artificial intelligence (AI) to achieve significantly improved performance. The growing effect of AI-based tools in protein science necessitates enhanced educational materials that improve AI literacy among established scientists seeking to deepen their expertise and new researchers entering the field. To address this need, we developed Deep Learning for Proteins: a series of 10 interactive notebook modules that introduce fundamental machine-learning concepts, guide users through training machine-learning models for protein-related tasks, and ultimately present cutting-edge protein structure prediction and design pipelines. By using only a web browser, learners can access state-of-the-art computational tools used by professional protein engineers that range from all-atom protein design to fine-tuning protein language models for biophysically relevant functional tasks. By increasing accessibility, this notebook series broadens participation in AI-driven protein research. The complete notebook series is publicly available at
https://github.com/Graylab/DL4Proteins-notebooks
.
This mini review traces the evolution of AI-driven methods in protein research, from early residue-contact prediction using coevolutionary information to transformative breakthroughs, the rise of protein language models (PLMs), and the emerging era of generative design and functional modeling.
Guodong Min, Huan Peng· Methods in molecular biology· 0 citations
Deep-Interact Studio is, to the authors' knowledge, the only such platform to combine fine-grained per-layer model customization with multi-model comparison and interpretability, offering a flexible and transparent alternative to fixed, single-purpose tools.
Dipayan Sarkar, K. Bardhan, Chiranjib Sarkar· bioRxiv· 0 citations
This Perspective explores recent methods and prospective ideas for developing hybrid AI-physics-based pipelines for protein and antibody de novo design. We argue that the highest-confidence candidates emerge where deep learning and first-principles models agree, a "sweet spot" that balances generative flexibility with thermodynamic realism. For example, although interface confidence scores such as ipTM, pDockQ2, or ipSAE are widely used to rank generated designs, we show that they are not well suited to rank similar sequences, which suggests the need to combine them with physics-based methods to improve design filtering and ranking. Furthermore, we describe a generalizable framework for implementing antibody design pipelines that combine AI with physics-based modeling and scoring methods and also showcase MadraX, a differentiable and AI-compatible implementation of the FoldX force field. In addition, we classify three tiers of AI-physics integration, from post hoc filtering to full embedding of differentiable physics inside deep learning models. Finally, we discuss the future of the protein design community and underline the need to support current initiatives for community wide blind assessments of the growing number of de novo design pipelines.
Gabriel Cia, Gabriele Orlando, D. Cianferoni et al.· FEBS Letters· 0 citations
Addressing and predicting ligand-binding sites in protein structures, as well as the prediction of reliable structures of proteins interacting with other proteins, will be pivotal for fully details of structural mechanisms and dynamics.
Pradeep Bk, Shi-Jie Chen, R. Dima et al.· Journal of Molecular Biology· 0 citations
Predicting protein stability, like changes in melting temperature (ΔTm) caused by mutations, is a critical task in therapeutic protein engineering and drug discovery. This is reflected by a growing solution space, including both AI-based sequence and structure based methods. This paper demonstrates that accurate ΔTm prediction does not require structural input features, but can achieve state-of-the-art results with a careful training design for large sequence-based protein language models. We combine an autoresearch-inspired setup search with controlled ablation studies and show that a well-tuned sequence-only ESM2-650M model [6] outperforms structure-informed methods in our benchmark, achieving the lowest error (MAE/RMSE) and competitive Pearson correlation without pH or structural inputs. We further show that choices such as loss function, pooling strategy, auxiliary supervision, and finetuning regime materially affect performance.
Daniel Siegismund, Mario Wieser, E. Natali et al.· bioRxiv· 0 citations