Skip to content

Deep Learning Foundation Models for Low-Data Regimes from Classical Molecular Descriptors

Aug 2026 · Journal of Chemical Information and Modeling · 0 citations · 53 references

TL;DR

This work proposes pretraining on low-noise, calculable molecular descriptors via supervised learning to obtain rich, highly transferable molecular representations and demonstrates this strategy with CheMeleon, a O(10M) parameter foundation model that enables directed message-passing neural networks to finally exceed the performance of classical methods in the low-data regime.

Abstract

Fast and accurate data-driven prediction of molecular properties is pivotal to scientific advancements across myriad chemical domains. Deep learning methods have recently garnered much attention, despite their inability to outperform classical machine learning methods when tested on practical, real-world benchmarks with limited training data. This study seeks to bridge this gap by introducing a new avenue for foundation model pretraining. We propose pretraining on low-noise, calculable molecular descriptors via supervised learning to obtain rich, highly transferable molecular representations. We demonstrate this strategy with CheMeleon, a O(10M) parameter foundation model that enables directed message-passing neural networks to finally exceed the performance of classical methods in the low-data regime. We evaluate on 58 benchmark data sets spanning a range of properties relevant to small-molecule drug discovery, sourced from the industry-led Polaris benchmarking initiative. Rigorous statistical comparisons show that CheMeleon outperforms classical baselines like Random Forest on molecular fingerprints and descriptors, as well as existing foundation models. We open-source the CheMeleon model and the pretraining framework to encourage adoption and extension of this pretraining strategy across chemical sciences.

View source

Similar papers

Preprint Aug 2026

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Monroe is presented, a new MFM with several innovations over the existing state of the art: increased scale allowing pre-training on over 81 million molecules from the PM6 quantum chemistry dataset; improved graph representation of stereochemistry; improved training losses including conformer denoising and embedding decorrelation; improved multi-task learning; and the use of a prior-data-fitted model (TabPFN) for downstream in-context prediction.

Blazej Banaszewski, Andrew W. Fitzgibbon · 0 citations
Aug 2026

Rapid Generative Discovery of High‐Energy Molecules in Low Data Regimes Using Minimal Computational Resources

This work presents a novel approach toward high‐energy molecules by combining long short‐term memory (LSTM) networks for molecular generation and attentive graph neural networks (GNN) for property predictions by combining fixed SHA‐256 embeddings with partially trainable representations.

Siddharth Verma, A. Alankar · 0 citations
Review Jul 2026

Self-Supervised Learning for Molecular Property Prediction: Methods, Multimodal Insights, and Benchmark Comparisons.

Computer-aided drug discovery has substantially accelerated modern pharmaceutical research, where accurate molecular property prediction plays a central role in identifying promising therapeutic candidates. Self-supervised learning (SSL), which exploits large-scale unlabeled molecular data to learn transferable representations, has recently emerged as a powerful paradigm well-aligned with the data characteristics of cheminformatics. Integrating chemical domain knowledge further enhances the ability of SSL models to capture structural, physicochemical, and functional properties of molecules. In this review, we provide a systematic overview of recent advances in SSL-based molecular property prediction. We summarize representative methodological developments and analyze how multimodal molecular representation learning─by integrating sequence, graph, three-dimensional structure, and textual information─can improve the quality and expressiveness of molecular representations. We further examine the synergistic relationship between multimodal modeling and SSL, highlighting how complementary modalities can enhance representation learning in low-label settings. To demonstrate the practical benefits of multimodal molecular properties, we compare their performance with conventional SSL models on two downstream benchmark tasks with distinct prediction objectives. Finally, we discuss key open challenges, including the scarcity of high-quality 3D molecular data, modality imbalance across data sets, and the limited interpretability of learned representations. We conclude by outlining promising research directions toward more robust, generalizable, and biologically meaningful frameworks for molecular property prediction.

Shuning Yang, Lei Deng · 0 citations
Open access Aug 2026

Synthetic data for more accurate deep learning models in molecular science: a test case of protein-ligand binding affinity prediction

This study shows that incorporating synthetic molecular dynamics data improves deep learning models for protein–ligand binding affinity prediction beyond static experimental structures, and highlights that dynamic synthetic datasets can enable deep learning models to outperform conventional methods such as MM-PBSA while remaining computationally efficient.

P. Agrawal, Prathit Chatterjee, U. Priyakumar · 0 citations
Open access Jul 2026

P2MAT: A machine learning (ML) driven software for Property Prediction of MATerial.

Accurate prediction of melting points for pure molecules remains a significant challenge in predictive chemistry, with implications across various scientific fields, including materials science, drug discovery, and separations chemistry. Traditional methods, such as group contribution (GC) techniques, have shown limited success due to the complex relationship between molecular structure and melting point. In this study, we present a data-driven machine learning (ML) approach to predict the melting points of organic compounds, leveraging both 2D and 3D molecular descriptors. Our results indicate that ML models can significantly improve melting-point predictions, providing a robust tool for the scientific community. Scientific contributionOur detailed analysis on melting point prediction, along with SHAP explainability, reveals the top influencing features for the prediction. The P2MAT application we developed as part of this study can predict both melting and boiling points from a SMILES string. P2MAT is available as an easy-to-install, user-friendly GUI for maximum outreach to the scientific community. Our benchmark analysis demonstrates the excellence of our method for predicting melting points.

Md Kamruzzaman, Alexander Landera, N. Menon et al. · 1 citation
Aug 2026

Dual-Attention Multimodal Framework for Molecular Property Prediction

A novel Dual-Attention Multimodal framework for Graphs and Sequence-based representations, so-called DAM-GS, which provides a promising solution for molecular property prediction with broad applications in drug discovery and computational molecular science.

Bay Van Nguyen, Vinh Truong, Ha Duong Thi Hong et al. · 0 citations