Skip to content
Open access

AI-driven drug discovery using transformer-based molecular representation learning

Jul 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 37 references
Medicine

TL;DR

A transformer-based molecular modeling framework for target-specific potency prediction, trained on curated BindingDB bioactivity data across Alzheimer‘s, diabetes, and cancer targets to deliver accurate pIC50 regression and binary activity classification.

Abstract

The vast majority of chemically plausible, drug-like molecules remain unexplored due to the combinatorial scale of chemical space and the limited throughput of experimental screening. This is complicated by the lack of data and the inability to extrapolate predictive models to chemotypes that are not represented well. We introduce a transformer-based molecular modeling framework for target-specific potency prediction, trained on curated BindingDB bioactivity data across Alzheimer‘s, diabetes, and cancer targets to deliver accurate pIC50 regression and binary activity classification. It uses curated bioactivity data from BindingDB to build target-specific datasets and uses a Byte Latent Transformer (BLT) that is trained directly on SMILES strings to predict changes in compound activity and potency based on quantitative structure–activity relationships. The transformer captures both syntactic and higher-level chemical features of SMILES representations and performs byte-level predictions using a latent model trained on the same molecular information. Potency predictions are executed inside an engine of chemically described molecular search engine, which executes stochasticity, SMILES-based amount mutations benefit by the envisaged activity, drug-like rules, and adaptive seeking heuristics to prevent local minima. Optimized candidate molecules and local optima are generated through guided SMILES mutations, preserving structural diversity around high-potency leads. The framework delivers highly accurate pIC50 regression (R2 0.95–0.98) and binary activity classification across Alzheimer's, diabetes, and cancer targets, enabling robust virtual screening and lead prioritization for diverse biological targets.

Read PDF

Similar papers

Open access Jul 2026

Smiles-based bioactivity prediction through molecular encoder selection and data augmentation.

Quantitative prediction of inhibitor potency can accelerate early-stage drug discovery. Recently, data-driven approaches have gained widespread interest in drug discovery, as evidenced by a growing number of benchmarking challenges and open competitions. In this context, we developed a machine learning-based methodology that can find the most effective way of predicting IC50 values against ASK1 from SMILES, for "Jump AI(.py) 2025: 3rd AI Drug Discovery Competition", hosted by the Korea Pharmaceutical and Bio-Pharma Manufacturers Association (KPBMA) on the Dacon platform. Applying our methodology achieved the highest overall predictive performance among all participating teams. Beyond this competition setting, we present a compact SMILES-based modeling workflow comprising (i) a pre-trained encoder, (ii) regression models, (iii) data augmentation, and (iv) hyperparameter tuning. We systematically compared molecular representations from sequence- and graph-based models, including ChemBERTa-2 and MolCLR. Across encoder-regressor combinations, ChemBERTa-77 M-MLM embeddings paired with support vector regression (SVR) yielded the strongest predictive performance. Embedding-level mix-up augmentation and SVR hyperparameter tuning further improved predictive performance. Our findings highlight that careful SMILES preprocessing and encoder selection have a critical influence on IC50 values and provide a reproducible benchmark for single-target bioactivity prediction, thus contributing to a more efficient drug discovery process. Scientific Contribution In this study, we propose a machine learning methodology for predicting the IC50 values of ASK1 inhibitors from SMILES representations, with a systematic comparison of molecular encoders and regression models. Our results show that the use of suitable encoder-regressor pairs together with embedding-level mix-up augmentation improves model generalizability without requiring SMILES-level augmentation. This strategy would be particularly useful for settings with imbalanced labels or limited data, and could be applied more broadly to IC50 prediction for other kinase inhibitors.

Ju Hyung Lee, S. Choi, Utku Ozbulak et al. · 0 citations
Open access Jul 2026

Transformer-based molecular fragment prediction using SMILES and DeepSMILES representations in a fragment-based drug discovery pipeline

Fragment-based drug discovery (FBDD) uses small molecular fragments as starting points for drug development, and machine-learning models that operate on molecular string representations are increasingly applied to fragment-related tasks. SMILES is the dominant such representation, but its paired ring closures and balanced parentheses introduce syntactic complexity that can affect model behavior. We present a controlled study of how molecular string representation influences transformer-based fragment recovery, using an FBDD-motivated label pipeline: reference fragments are derived from known drugs via RECAP fragmentation and docking-based ranking, and the model is scored on recovering them. We introduce DeepBERTa, a ChemBERTa-derived transformer pretrained on DeepSMILES, and evaluate it against a matched SMILES baseline under identical architecture, data splits, and optimization (approximately 34,000 drug–fragment pairs). DeepSMILES produces syntactically valid predictions more often than SMILES (54.2% vs. 43.5%) and a higher full test-set mean Tanimoto similarity (0.36 vs. 0.29). A per-sample selection between the two representations raises mean Tanimoto to 0.43, and a deployable variant that selects on model confidence rather than the reference recovers most of this gain. Among molecules where one representation strictly wins, DeepSMILES wins more often than SMILES (27.8% vs. 16.3% of test molecules). These results show that string representation meaningfully influences transformer fragment recovery and that representation-aware selection outperforms either representation alone.

Aayush Kothari, Amisha Gupta, Nisarg Shah et al. · 0 citations
Open access Aug 2026

Learning from human and chemical languages to predict biological function

PubCheF-1, a deep learning model that predicts literature-derived biological function directly from chemical structure, establishes that machine learning-based prediction of biological function derived from the language of scientific literature allows the identification of bioactive molecules at high hit rates, thereby accelerating therapeutic discovery.

Clayton W. Kosonocky, Nikol Kadeřábková, Kangsan Kim et al. · 0 citations
Aug 2026

Dual-Attention Multimodal Framework for Molecular Property Prediction

A novel Dual-Attention Multimodal framework for Graphs and Sequence-based representations, so-called DAM-GS, which provides a promising solution for molecular property prediction with broad applications in drug discovery and computational molecular science.

Bay Van Nguyen, Vinh Truong, Ha Duong Thi Hong et al. · 0 citations
Conference Jul 2026

AI-Driven Virtual Screening of Naturalcompounds Against SARS-CoV-2 Using Embedding-Based Drug-Protein Interaction Prediction

The imperative necessity for rapid discovery of antiviral agents against emerging viral diseases, such as COVID-19 caused by SARS-CoV-2, has emphasised the limitations of conventional drug discovery, with its deliberate pace, high costs, and high failure rate. In this study, we present a high-throughput “AI-driven virtual screening pipeline” that combines cutting-edge molecular embeddings using transformers for compounds and language model-based protein sequences for targets, and combines these using a gradient-boosted decision tree-based model (XGBoost) trained on ChEMBL bioactivity data, with excellent performance for binary prediction (ROC-AUC 0.8408, Accuracy 0.76) comparable to top-performing methods. When applied to the large library of natural products from COCONUT database, our model correctly predicted top-ranked compounds, some of which were shortlisted for molecular docking against the SARS-CoV-2 spike receptor-binding domain (RBD), the key interface for ACE2 binding, in order to test the validity of proposed pipeline and the results revealed several promising compounds with excellent binding energies and multiple interactions with hotspot residues such as K417, Y453, Q493, G496, Q498, N501, Y505 and F486. The potential of cost-effective, synergistic integration of scalable deep learning of representations, interpretable machine learning, and physics-based refinement of structure is significant for accelerating natural product-based therapeutics for coronaviruses and possible other viral threats, with the potential to expand to ensemble approaches, active learning, and variant-based targets for enhanced efficacy.

Hayat Ullah, Khan Ziaullah, Md Ariful Islam Mozumder et al. · 0 citations
Open access Aug 2026

From Descriptor Learning to Binding Stability: An Explainable Machine Learning Pipeline for EGFR Double-Mutant Inhibitor Discovery

An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.

Jurica Novak · 0 citations