Skip to content

Scalable and Generalizable Analog Design via Learning Medicinal Chemistry Intuition from Matched Molecular Pair Transformations.

Jul 2026 · Journal of Chemical Information and Modeling · Vol 66 15, pp. 8908-8922 · 0 citations · 39 references
Medicine

TL;DR

This work introduces a novel approach to transform the concept of "how to design chemical analogs like a medicinal chemist" into a primary training objective for generative models by focusing on matched molecular pair transformations (MMPTs) as the fundamental unit of chemical modifications.

Abstract

Chemical analog design in the hit-to-lead and lead optimization stages of drug discovery relies on systematic structural modifications, often guided by medicinal chemistry intuition. Although a matched molecular pair (MMP) provides an interpretable framework to capture intuition, models trained on individual MMP instances face significant limitations, such as bias toward frequent transformations in historical data. This work introduces a novel approach to transform the concept of "how to design chemical analogs like a medicinal chemist" into a primary training objective for generative models by focusing on matched molecular pair transformations (MMPTs) as the fundamental unit of chemical modifications. This allows for a more generalizable and context-independent representation of medicinal chemistry intuition, enabling the application of the same transformation priors across different projects, regardless of the target or indication. Using a consistently curated ChEMBL-derived data set, we compared a transformation-centric foundation model (MMPT-FM) with multiple MMP-based generative formulations trained on the same underlying data. Furthermore, performance is assessed through challenging within-patent and cross-patent real-world test cases derived from drug discovery patents. The MMPT-FM model achieves comparable or improved recall metrics across all test cases, demonstrating particularly strong performance for low-frequency and previously unseen transformations. This work not only shifts the paradigm of learning and utilizes medicinal chemistry intuition in an efficient and scalable manner in the AI for drug discovery era but also establishes a significant competitive advantage by enabling the drug discovery industry to encode decades of collective medicinal chemistry knowledge─both public and proprietary─into a scalable foundation generative model that could help researchers design chemical analogs.

View source

Similar papers

Review Aug 2026

The Evolution of Generative Chemistry in Medicinal Chemistry

Generative chemistry is an emerging discipline that utilizes generative artificial intelligence (AI) models for the automated de novo design of small molecules. By learning patterns from existing chemical data, these models can generate novel structures with desired properties, thereby accelerating drug discovery. However, a significant gap remains between the potential of AI and its successful implementation in practical pharmaceutical applications. This review covers various infrastructural aspects of generative chemistry, including molecular representation, databases, and diverse model architectures such as generative adversarial networks, variational autoencoders, and diffusion models. Furthermore, key challenges associated with data quality, model selection, and synthesis feasibility are critically discussed. The review highlights that generative chemistry has evolved beyond simple structure generation to encompass the entire molecular design pipeline, including automated synthesis planning, retrosynthesis prediction, and multi‐objective optimization. Additionally, the selection of the most suitable model depends on specific objectives and the quality and diversity of the dataset, rather than a single superior architecture. Overall, a critical perspective is provided on how generative models are shaping the future of rational and reliable drug design.

Rania Ehab Koshty, Manar Ahmed Shehata, Ahmed M. Gab Allah et al. · 0 citations
Open access Jul 2026

ScrambleBench: a workflow for comparative assessment of structure-based de novo generative models.

ScrambleBench provides a holistic medicinal chemistry-oriented framework that identifies methodological strengths, limitations, and opportunities for future model development and highlights the importance of evaluating chemical diversity explicitly and using the recently proposed metrics such as Hamiltonian Diversity (HamDiv) which assess both quantity and dissimilarity of a molecular set.

Veincent Yap, Pan Xu, Frankie S. Mak et al. · 0 citations
Preprint Jul 2026

Vilya-1: An all-atom foundation model for macrocycle structure prediction and design

Vilya-1 is introduced, a deep learning model that addresses two central challenges in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries and predicting key developability properties such as membrane permeability.

Vilya Research Pascal Sturmfels, M. Salem, Naozumi Hiranuma et al. · 1 citation
Preprint Aug 2026

Packora: Systematic Design for Generative Molecular Crystal Structure Prediction

Molecular crystal structure prediction (CSP) is important in pharmaceuticals, agrochemicals, and organic electronics, where subtle differences in molecular conformation and packing can strongly affect material properties. We present Packora, a flow-based generative model for molecular CSP that jointly predicts atomic coordinates and the lattice from molecular graphs. Packora supports multi-component and organometallic crystals and can condition on any subset of molecular conformers, stereochemical labels, and space-group information within a single model. Inspired by the CCDC CSP blind test, we evaluate generation and ranking separately, using generation to isolate generator quality and ranking to measure end-to-end performance under a common relaxation and ranking pipeline. We also systematically study architecture, training, conditioning, inference, and scaling, identifying an effective design based on cacheable pairwise reasoning, training objective and numerical solver choices, conditioning dropout, and balanced scaling of pairwise and single representations. Packora outperforms the baselines on both structure generation and ranking benchmarks, achieving the best matched-budget coverage across all six generation benchmarks, as well as higher experimental-form recovery, lower experimental-form ranks, and faster convergence in ranking.

Nayoung Kim, Kiyoung Seong, Sungsoo Ahn · 0 citations
Preprint Jul 2026

Sample Efficient Generative Optimization for Molecular Design

Molecular optimization in drug discovery, materials design, and catalysis requires searching vast chemical spaces under tight evaluation budgets, since high-fidelity oracles and experimental measurements are costly. The practical impact of an optimization method therefore hinges on its sample efficiency: how few evaluations it needs to find strong candidates. We introduce Sample Efficient Generative Optimization (SEGO), a framework for Bayesian optimization on adaptively generated molecules. In SEGO, a probabilistic surrogate model forms a hypothesis about where hits lie in chemical space, a generative model is steered to propose candidates in that region, the most promising candidate is selected via an acquisition function, and the resulting oracle call is used both to sharpen the surrogate and to anchor the generator in real reward. SEGO attains state-of-the-art performance on the practical molecular optimization (PMO) benchmark using only one tenth of the oracle calls consumed by other methods, and on a multiparameter docking task it reaches ten hits in roughly half the oracle calls of existing approaches. These gains move molecular optimization closer to campaigns driven by direct experimental feedback.

S. Kopf, Cristina Nevado, P. Schwaller · 0 citations
Preprint Aug 2026

Synthesizing like a chemist: an iterative, feedback-driven loop for materials discovery

Most computationally predicted materials are never synthesized because conventional synthesis optimization is slow, expertise-dependent, and iterative. Here we present a closed-loop framework that automates this expert workflow by placing human tacit knowledge in the loop through a large language model (LLM) that distills synthesis knowledge from the literature, high-throughput hyperspectral imaging for rapid film evaluation, and multi-objective Bayesian optimization guided by experimental feedback. In a paired optimization campaign, LLM-assisted initialization produced more Pareto-optimal samples and higher hypervolume than a Latin hypercube sampling baseline at matched trial counts, and this advantage persisted throughout iterative optimization. We demonstrate the framework by synthesizing the previously unreported perovskite-inspired compound Rb3BiI6 as thin films and validating the optimized films by optical bandgap analysis and X-ray diffraction. The framework transforms synthesis prediction from single-shot recommendation to iterative learning, providing a generalizable strategy to accelerate automated and fully autonomous experimental materials discovery.

Fang Sheng, Steven B. Torrisi, Amanda A. Volk et al. · 0 citations