Skip to content
Open access

Synthetic data for more accurate deep learning models in molecular science: a test case of protein-ligand binding affinity prediction

Aug 2026 · Journal of Cheminformatics · 0 citations

TL;DR

This study shows that incorporating synthetic molecular dynamics data improves deep learning models for protein–ligand binding affinity prediction beyond static experimental structures, and highlights that dynamic synthetic datasets can enable deep learning models to outperform conventional methods such as MM-PBSA while remaining computationally efficient.

Abstract

Deep learning models are data-hungry, and synthetic (artificial) data has been shown to be invaluable when data availability is low. While this has been demonstrated in certain technology areas, adopting such an approach is new in machine learning (ML) applications in chemistry, except for some pre-training tasks. In drug discovery, predicting binding energy between proteins and ligands is crucial. Many ML-based studies have been proposed to predict protein-ligand binding affinity using existing experimental data. However, these models suffer from inherent biases. Recent efforts have produced PLAS-20k, a synthetic dataset of multiple protein-ligand complex (PLC) conformations generated using molecular dynamics (MD) simulations as a viable option to complement existing experimental data and improve binding affinity prediction. For the binding affinity prediction task, we employ Pafnucy, a deep convolutional neural network, and propose using multiple structures for each PLC from PLAS-20k for training. We compare four different statistical and ML-based result-aggregation techniques. This work demonstrates the utility of dynamic datasets in enhancing binding affinity predictions, laying the foundation for future improvements in predicting similar protein properties using synthetic datasets and more sophisticated models and methods. We propose that physics-based synthetic datasets can significantly help develop more accurate data-driven methods. Scientific contribution This study shows that incorporating synthetic molecular dynamics data improves deep learning models for protein–ligand binding affinity prediction beyond static experimental structures. By systematically evaluating frame selection and prediction aggregation strategies, we demonstrate that training on diverse conformational snapshots significantly enhances generalization and accuracy. Our results highlight that dynamic synthetic datasets can enable deep learning models to outperform conventional methods such as MM-PBSA while remaining computationally efficient.

Read PDF

Similar papers

Jul 2026

Bridging between Structure-Based and Data-Driven Affinity Prediction.

This work introduces a method to smoothly transition from physics-based to knowledge-based predictions based on the uncertainty of each model and shows that combining structure-based and ML models significantly improves the prediction accuracy if training data is limited, whereas the weighting smoothly shifts from docking to ML as more data is acquired.

Ažbeta Kubincová, David L. Mobley · 1 citation
Open access Jul 2026

A Preparation-Free Mixture-of-Experts Framework for Protein-Ligand Affinity Prediction

The resulting model, HydrAffinity, is an interaction-free, dynamic sparse model that uses pre-trained encoders and MoE for parameter-efficient learning and outperforms all interaction-free methods and matches state-of-the-art interaction-based methods on CASF-2016.

Huiming Bao, Shouliang Dong · 0 citations
Review Aug 2026

Recent Advances in Deep Learning-Based Drug-Target Binding Affinity Prediction

It is indicated that although many methods report strong performance on standard benchmarks, their effectiveness is often influenced by dataset bias and limited evaluation settings, and most methods exhibit reduced performance in cold-start scenarios, highlighting challenges in generalization.

Jafin Khan, Md Hossain Shuvo · 0 citations
Review Open access Aug 2026

Geometric Deep Learning‐Based Drug Design Models for Small‐Molecule Drug Discovery

Deep neural network (DNN)‐based in silico models show great promise in predicting the properties and bioactivities of novel compounds, including small molecules. Among traditional approaches, structure‐based drug design (SBDD) remains a fundamental approach for drug discovery using molecular docking, scoring functions, and molecular dynamics simulations. However, these approaches are often constrained by limited flexibility, resolution, and generalizability. Geometric deep learning (GDL) offers a transformative alternative by enabling models to learn directly from non‐Euclidean molecular representations, such as graphs, point clouds, and meshes, capturing critical 3D spatial relationships inherent to protein–ligand interactions. This review highlights the theoretical underpinnings and practical applications of GDL in small‐molecule drug discovery, focusing on tasks including binding affinity prediction, virtual screening, de novo molecule generation, pose prediction, ADMET profiling, and protein flexibility modeling. We explore key GDL architectures, graph neural networks, SE(3)‐equivariant networks, 3D convolutional neural networks, point cloud models, and geometric transformers, and assess their performance across various drug discovery benchmarks. The integration of geometry‐aware AI models with experimental and computational workflows was also highlighted for its potential to streamline hit‐to‐lead optimization and advance rational drug design. Despite remarkable progress, the field faces challenges including limited high‐quality 3D structural datasets, protein flexibility representation, and the interpretability of deep models. Addressing these issues through hybrid modeling approaches, multi‐resolution learning, and self‐supervised training could further elevate GDL's impact. Ultimately, GDL stands at the frontier of AI‐enhanced pharmaceutical innovation, offering unprecedented precision, efficiency, and insight in the pursuit of next‐generation therapeutics.

A. Srivastav, Unnati Modi, Rahul Kumar et al. · 0 citations
Aug 2026

Deep Learning Foundation Models for Low-Data Regimes from Classical Molecular Descriptors

This work proposes pretraining on low-noise, calculable molecular descriptors via supervised learning to obtain rich, highly transferable molecular representations and demonstrates this strategy with CheMeleon, a O(10M) parameter foundation model that enables directed message-passing neural networks to finally exceed the performance of classical methods in the low-data regime.

Jackson W. Burns, Akshat Shirish Zalte, C. Abreu et al. · 0 citations
Preprint Aug 2026

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Monroe is presented, a new MFM with several innovations over the existing state of the art: increased scale allowing pre-training on over 81 million molecules from the PM6 quantum chemistry dataset; improved graph representation of stereochemistry; improved training losses including conformer denoising and embedding decorrelation; improved multi-task learning; and the use of a prior-data-fitted model (TabPFN) for downstream in-context prediction.

Blazej Banaszewski, Andrew W. Fitzgibbon · 0 citations