Skip to content
Open access

Learning from human and chemical languages to predict biological function

Aug 2026 · bioRxiv · 0 citations
Biology

TL;DR

PubCheF-1, a deep learning model that predicts literature-derived biological function directly from chemical structure, establishes that machine learning-based prediction of biological function derived from the language of scientific literature allows the identification of bioactive molecules at high hit rates, thereby accelerating therapeutic discovery.

Abstract

Understanding how molecular structure encodes biological function remains a grand challenge in drug discovery. Here, we present PubCheF-1, a deep learning model that predicts literature-derived biological function directly from chemical structure. PubCheF-1 was trained on a dataset linking molecules to labels derived from the scientific articles in which they appear, a strategy that connects disparate compounds through the language used to describe their functionalities. When tasked with identifying inhibitors of β-lactamases, including enzymes considered largely refractory to inhibition, PubCheF-1 predicted structurally distinct compounds that collectively have activity against all β-lactamase classes. Furthermore, hit compounds directly bind the enzyme active site, restore antibiotic efficacy in multidrug-resistant high-priority pathogens, and demonstrate potent activity in animal infection models. Together, these findings establish that machine learning-based prediction of biological function derived from the language of scientific literature allows the identification of bioactive molecules at high hit rates, thereby accelerating therapeutic discovery.

Read PDF

Similar papers

Review 2026

AI-Driven Protein Research: From Prediction to Design.

This mini review traces the evolution of AI-driven methods in protein research, from early residue-contact prediction using coevolutionary information to transformative breakthroughs, the rise of protein language models (PLMs), and the emerging era of generative design and functional modeling.

Guodong Min, Huan Peng · 0 citations
Open access Jul 2026

High-resolution dissection of concept acquisition in different families of protein language models

A high-resolution layer-by-layer interpretability analysis of 8 models from the ESM2 and AMPLIFY families on 22 concepts from human proteome annotations found that these models encode concepts of increasing levels of complexity along their depth: basic physicochemical properties and linear motifs are best captured by early-layer embeddings, secondary structure from subsequent layers, and domain-level semantics from middle layers.

Shawn T. Whitfield, Tom Marty, Robert M. Vernon et al. · 0 citations
Preprint Jul 2026

Accurate structural modeling of chemically diverse molecular interfaces with Vilya-2

Vilya-2 is the structure-prediction oracle that de novo peptide design pipelines require--establishing the all-atom approach as a general foundation for the design and evaluation of de novo peptide therapeutics.

Vilya Research Pascal Sturmfels, Naozumi Hiranuma, M. Salem et al. · 0 citations
Preprint Jul 2026

DrugGen 2: A disease-aware language model for enhancing drug discovery

Current computational approaches for drug design typically focus on generating molecules conditioned on specific targets or general molecular properties, often neglecting the influence of disease context on target behavior and therapeutic outcomes. To address this gap, we introduce DrugGen-2, a novel generative model that designs small molecules conditioned on both disease ontology and target protein sequences. DrugGen-2 was developed by fine-tuning a pre-trained GPT-2 model on a curated dataset of approved drugs linked to their diseases and targets, using a two-step strategy of supervised fine-tuning followed by reinforcement learning via group relative policy optimization (GRPO). This process was guided by reward functions optimizing for chemical validity, novelty, diversity, and high predicted binding affinity. When evaluated on five protein targets relevant to diabetic nephropathy, DrugGen-2 significantly outperformed baseline models (DrugGPT and DrugGen). It demonstrated a superior capacity to generate unique molecules, exhibited greater structural similarity to approved drugs, and achieved improved predicted binding affinities across all targets. Molecular docking analyses further supported these findings, identifying candidate ligands with strong binding potential, including compounds with predicted affinities (-9.917, -9.485, and -9.367) exceeding those of reference drugs such as enalapril for angiotensin-converting enzyme (-8.283). By integrating disease-specific context into molecular generation, DrugGen-2 advances AI-assisted drug discovery, offering a powerful tool for de novo design and drug repurposing that accounts for the complex interplay between diseases and molecular targets.

Ali Motahharynia, Mohammadreza Ghaffarzadeh-Esfahani, Mahsa Sheikholeslami et al. · 0 citations
Review Jul 2026

Self-Supervised Learning for Molecular Property Prediction: Methods, Multimodal Insights, and Benchmark Comparisons.

Computer-aided drug discovery has substantially accelerated modern pharmaceutical research, where accurate molecular property prediction plays a central role in identifying promising therapeutic candidates. Self-supervised learning (SSL), which exploits large-scale unlabeled molecular data to learn transferable representations, has recently emerged as a powerful paradigm well-aligned with the data characteristics of cheminformatics. Integrating chemical domain knowledge further enhances the ability of SSL models to capture structural, physicochemical, and functional properties of molecules. In this review, we provide a systematic overview of recent advances in SSL-based molecular property prediction. We summarize representative methodological developments and analyze how multimodal molecular representation learning─by integrating sequence, graph, three-dimensional structure, and textual information─can improve the quality and expressiveness of molecular representations. We further examine the synergistic relationship between multimodal modeling and SSL, highlighting how complementary modalities can enhance representation learning in low-label settings. To demonstrate the practical benefits of multimodal molecular properties, we compare their performance with conventional SSL models on two downstream benchmark tasks with distinct prediction objectives. Finally, we discuss key open challenges, including the scarcity of high-quality 3D molecular data, modality imbalance across data sets, and the limited interpretability of learned representations. We conclude by outlining promising research directions toward more robust, generalizable, and biologically meaningful frameworks for molecular property prediction.

Shuning Yang, Lei Deng · 0 citations
Open access Sep 2025

LINKER: Learning Interactions between Functional Groups and Residues with Chemical Knowledge‑Enhanced Reasoning and Explainability

LINKER is the first sequence-based model to predict residue-functional group interactions according to biologically defined interaction types, using only a protein sequence and the SMILES representation of the ligand, and requires only sequence-level input at inference.

Phuc Pham, Viet Thanh Duy Nguyen, Truong-Son Hy · 1 citation