Aug 2026· Nature Computational Science· 0 citations· 49 references
Medicine
TL;DR
Self-supervised pretraining substantially improves TS prediction for previously unseen systems, lowering the median root-mean-square deviation of TS geometries on Transition1x-TMC reactions and reducing fine-tuning data requirements, enabling reliable performance even in low-data regimes.
Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we introduce Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions. We compile CREED-CCV-2+USPTO-XL, an ultra-large-scale dataset of ~45.6 million verified reactions to train the C3LM (Chemistry Constraint-Consistent Language Model). By integrating fine-tuning with ChemCensor-based and novelty-oriented rewards, our model achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark. Further analysis of reaction uniqueness shows that LLMs and conventional models explore complementary reaction spaces, motivating ensemble-based retrosynthesis systems. Overall, our results establish Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.
B. Zagribelnyy, Ivan D. Ilin, N. Bondarev et al.· 0 citations
The Buchwald-Hartwig cross-coupling is a cornerstone of modern pharmaceutical synthesis, yet predictive modeling of its outcomes remains constrained by data quality and chemical space coverage. Electronic laboratory notebooks contain heterogeneous, noisy records, while open-source high-throughput experimentation (HTE) datasets are fragmented and narrow in scope, leading to poor model performance on unseen substrates and conditions. Here we introduce a framework that systematically standardizes and integrates multiple reaction datasets into a high-quality, unique-structure-per-entity dataset, coupled with active learning to strategically expand chemical space. By merging published Buchwald-Hartwig HTE data with new experimental results, we achieve a model with predictive power across novel substrates and conditions, delivering improved out-of-distribution predictions compared with previous approaches. Crucially, model-guided reagent recommendations were validated experimentally, confirming the framework's utility to uncover unexplored reactivity. This work establishes a blueprint for robust machine learning in synthetic chemistry and enables preemptive in silico reagent screening to accelerate pharmaceutical discovery.
Paulo Neves, Bo Hao, Santeri Aikonen et al.· Nature Computational Science· 0 citations
Reaction yield prediction is a longstanding challenge in synthetic chemistry, with broad implications for route planning, scalability, and high-throughput experimentation (HTE). While recent machine learning (ML) approaches have demonstrated promise in modeling reactivity, they often use complex descriptors or deep architectures that are computationally expensive and limit interpretability and scalability. Here, we assess how much information is stored in simpler descriptors and whether model accuracy is improved by increasing the complexity of the descriptors. Using classical ML models trained on descriptors with different complexity levels, we benchmark predictive performance on four publicly available HTE data sets covering three diverse reaction data sets: Buchwald–Hartwig (BH) amination, Suzuki–Miyaura (SM) coupling, and the silicon–amine protocol (SLAP). Our evaluation furthermore discusses (1) generalization via component-wise data splitting, (2) robustness through external validation across data sets, and (3) performance across asymmetric yield distributions characteristic of HTE data. Contrary to conventional expectations, we find that simpler models with interpretable features can achieve competitive performance under rigorous validation protocols. Based on our findings, we formulate good practices for future studies in this area. For example, comparison to low-cost baseline models should become a requirement for future ML studies for reaction-yield prediction.
Idil Ismail, Gregory A Landrum, Sereina Riniker· Journal of the American Chem...· 0 citations
Forecasting the outcomes of transition-metal-catalyzed reactions is notoriously complex due to the interplay of diverse physical and chemical variables. A persistent computational bottleneck has been effectively merging broad electronic descriptors with the localized, three-dimensional geometry of the reactive site. To bridge this representation gap, we present ChemFusion, a hybrid neural network that fuses conventional electronic features with explicit 3D atomic coordinates. Using a cross-attention mechanism, the model enables global electronic states to dynamically attend to specific spatial constraints within un-pooled molecular point clouds. When benchmarked against a diverse library of cross-couplings, this approach delivers exceptional predictive performance, decisively surpassing traditional single-modality frameworks. Importantly, extracting the attention matrices reveals that the architecture autonomously learns to identify and penalize restrictive steric hindrances. This provides a physically grounded interpretability, demonstrating that spatially aware networks can navigate complex reaction sterics that standard statistical models typically miss.
The discovery of highly active polyethylene (PE) catalysts demands a systematic understanding of structure-condition-activity relationships in a vast chemical space. In this Letter, we present a data-driven framework combining explainable machine learning (ML) with large-scale virtual library generation. From a curated data set of 507 catalysts (bis(phenoxyimine) and bis(imino)pyridine ligands, seven metals), a gradient boosting regression (GBR) model achieves a test R2 of 0.91, outperforming convolutional and graph neural networks. SHAP analysis identifies topological (Chi2v), electronic (EState_VSA), and hydrophobic (SlogP_VSA) descriptors as governing activity and reveals a classical volcano-type temperature dependence, fundamentally governed by the Sabatier principle. A virtual library of 665 685 structures, constructed via combinatorial fragment assembly, extends the known chemical space substantially. High-throughput screening, coupled with SCscore filtering, yields 1090 synthetically accessible candidates with predicted activities exceeding 2 × 107 g mol-1 h-1. Substructure analysis uncovers metal-dependent design rules, in which early transition metals favor electron-deficient aromatics while late metals profit from moderately sized alkyls. This work establishes a practical route from experimental data to actionable catalyst designs.
Xuefeng Li, Haoke Qiu, Hanwen Pei et al.· Journal of Physical Chemistr...· 0 citations
This work not only establishes a pioneering paradigm for interpretable ML-driven force field refinement but also provides the first feature engineering solution incorporating chemical, physical, and structural information specifically designed for the machine learning of energetic molecular crystals.
Qi He, Pengju Wang, Xudong He et al.· Molecules· 0 citations