Skip to content

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Aug 2026

From VAEs to Diffusion and LLMs: Modern Generative Models for Molecular Discovery

Generative models are emerging as a key technology for accelerating molecular discovery in drug design, materials science, and catalysis by enabling efficient exploration of the vast chemical space of possible molecules. Recent advances in deep generative modeling—including variational autoencoders (VAEs), diffusion models, flow matching methods, and autoregressive transformer-based approaches—have produced a diverse toolkit for generating molecular structures and optimizing their properties. However, these paradigms are often studied independently, leaving many machine learning researchers without a clear understanding of their connections, strengths, and limitations in molecular applications. This tutorial provides a unified introduction to modern generative modeling approaches for molecular generation, covering their theoretical foundations, algorithmic design, and practical considerations for molecular representations such as 1D SMILES strings, 2D molecular graphs, and 3D structures. While the tutorial primarily focuses on generative models for de novo molecular design, we also briefly discuss how similar modeling paradigms extend to reaction prediction and retrosynthesis. By presenting these models within a common framework, the tutorial aims to equip ML researchers and AI-for-science practitioners with a clear conceptual map of the generative modeling landscape for molecular discovery and identify emerging research opportunities in this rapidly evolving area.

Kehan Guo, Yili Shen, Jeeyhun Hwang et al. · 0 citations
Preprint Aug 2026

MolecularCanvas: LLM-assisted Small-Molecule Drug Discovery via Structure-Guided Constraints

Small-molecule drug discovery relies on iterative molecular optimization, where chemists repeatedly modify candidate compounds to balance multiple competing properties such as efficacy, toxicity, and solubility. Recent advances in generative AI (GenAI) have shown promise in accelerating this process by automatically proposing new molecular structures or targeted modifications. However, existing GenAI-based molecular design tools remain poorly aligned with experts'real-world workflows. Specifically, they offer limited support for specifying structure-level modification intents on molecules, provide insufficient transparency into model-generated modifications, and lack integrated support for downstream property evaluation with external computational tools. To address these challenges, we introduce MolecularCanvas, an interactive system that enables users to iteratively construct an optimization context by integrating high-level goals, structure-level annotations, property constraints, and reference-based preferences. This context guides the generation of candidate molecules across diverse molecular structures. MolecularCanvas further enhances transparency by providing evidence for AI-generated suggestions and streamlines molecular evaluation by integrating commonly used computational tools for property assessment into a unified interface. Finally, a user study with 12 participants demonstrates the usefulness and effectiveness of MolecularCanvas in helping users optimize candidate molecules.

Haoyu Dong, Rui Sheng, Shuhao Zhang et al. · 0 citations
Open access Jul 2026

Machine learning-accelerated screening of hydroquinone analogs for proton-coupled electron transfer

Proton-coupled electron transfer (PCET) mediated by hydroquinone and related molecules is key to natural and artificial energy conversion. The reactivity of these molecules depends on their bond dissociation free energy (BDFE), but studying the relationship between structure and thermochemistry across this chemical space has been limited by challenging experimental setup and high computational expense. Here, we present the first use of the AIMNet2 neural network potential to calculate average BDFE (BDFEavg) values for the 2H+/2e− dehydrogenation of about 200 000 hydroquinone-like compounds, including vicinal diamines, diols, and dithiols. Benchmarking against DFT calculations for 168 substituted ortho-phenylenediamines (opda) shows good agreement (R2 ∼ 0.84). Our analysis finds that the BDFEavg of diamines ranges from 50 to 80 kcal mol−1 and can be systematically tuned by modifying the backbone and N-substitution: electron-withdrawing groups raise BDFEavg by up to 15 kcal mol−1, while lower aromaticity in furan and thiophene backbones decreases BDFEavg by approximately 10 kcal mol−1 compared to the phenyl systems (∼65 kcal mol−1). Validation through cyclic voltammetry and reactivity studies with quinone oxidants for selected compounds supports the computational results. This extensive thermochemical database and a web-based prediction tool developed as a result of this work will offer valuable resources for designing PCET reagents for catalysis, energy storage, and biomedical uses.

Rajdeep Sarma, Yiwen Wang, David D Hebert et al. · 0 citations