Skip to content

Rapid Generative Discovery of High‐Energy Molecules in Low Data Regimes Using Minimal Computational Resources

Aug 2026 · Propellants, explosives, pyrotechnics · 0 citations · 43 references

TL;DR

This work presents a novel approach toward high‐energy molecules by combining long short‐term memory (LSTM) networks for molecular generation and attentive graph neural networks (GNN) for property predictions by combining fixed SHA‐256 embeddings with partially trainable representations.

Abstract

High‐energy materials (HEMs) are critical for propulsion and defense domains, yet their discovery remains constrained by experimental data and restricted access to testing facilities. This work presents a novel approach toward high‐energy molecules by combining long short‐term memory (LSTM) networks for molecular generation and attentive graph neural networks (GNN) for property predictions. We propose a transformative embedding space construction strategy that integrates fixed SHA‐256 embeddings with partially trainable representations. Unlike conventional regularization techniques, this changes the representational basis itself, reshaping the molecular input space before learning begins. Without recourse to pretraining, one of the proposed models for low data regime achieves 74.3% validity and 47.1% novelty. The generated library exhibits a mean Tanimoto coefficient of 0.214 relative to the training set signifying the ability of framework to generate a diverse chemical space. We identified 37 promising candidates for further research, all exhibiting a predicted detonation velocity greater than 9 .

View source

Similar papers

Aug 2026

Deep Learning Foundation Models for Low-Data Regimes from Classical Molecular Descriptors

This work proposes pretraining on low-noise, calculable molecular descriptors via supervised learning to obtain rich, highly transferable molecular representations and demonstrates this strategy with CheMeleon, a O(10M) parameter foundation model that enables directed message-passing neural networks to finally exceed the performance of classical methods in the low-data regime.

Jackson W. Burns, Akshat Shirish Zalte, C. Abreu et al. · 0 citations
Open access Aug 2026

Machine Learning Unveils Isolated‐Surrounded Pt Motifs in High‐Entropy Alloys for Superior Low‐Temperature Ammonia Oxidation

The sluggish kinetics of the ammonia oxidation reaction constitute a critical bottleneck in the development of low‐temperature direct ammonia fuel cells. High‐entropy alloys (HEAs), owing to their diverse active sites, have emerged as promising catalysts. However, their vast compositional space makes traditional quantum chemical screening prohibitively expensive. In this study, we provide a computational proof of concept showing that the classical d ‐band center theory fails to predict the ammonia oxidation activity of quinary HEAs and exhibits a negligible correlation with the energy barrier of the rate‐determining step. To address this limitation, we employed the Sure Independence Screening and Sparsifying Operator (SISSO) method to construct a transparent and interpretable symbolic descriptor, achieving excellent predictive accuracy ( R 2  = 0.981) within the range of the training data. This descriptor extends beyond simple single‐electron parameters by integrating the synergistic effects of electron‐donating ability, lattice stiffness, and local electronegativity perturbations. This data‐driven approach reduces computational costs by several orders of magnitude relative to exhaustive density functional theory (DFT) screening and identifies an “isolated‐surrounded” geometric configuration as a highly active site that significantly enhances intrinsic catalytic activity from a thermodynamic perspective. Crucially, the predicted motif should be interpreted within the hydrazine‐mediated thermodynamic framework used here, and its thermodynamic superiority may shift if alternative kinetic pathways dominate under operating conditions. This structural motif promotes efficient NN coupling while suppressing site poisoning. Overall, this study provides a coordination chemistry‐based blueprint for the rational design of next‐generation catalysts with reduced platinum‐group‐metal content and offers a theoretical framework for future experimental validation.

Shangfeng Jiang, Ting Tao, Kexiang Guo et al. · 0 citations
Jul 2026

Data-Driven Exploration of the Polyethylene Catalyst Chemical Space via Machine Learning.

A data-driven framework combining explainable machine learning (ML) with large-scale virtual library generation with large-scale virtual library generation is presented, establishing a practical route from experimental data to actionable catalyst designs.

Xuefeng Li, Haoke Qiu, Hanwen Pei et al. · 0 citations
Aug 2026

Interpretable Machine Learning Framework Deciphers the Role of Local Environment of High‐Entropy Intermetallic Compounds for Alkaline Hydrogen Evolution Reaction

The application of high‐entropy intermetallic (HEI) compounds in the field of catalysis has attracted widespread attention, but their huge material space seriously hinders experimental exploration. Herein, we for the first time reported the efficient design, screening, and prediction of a great deal of high‐performance HER catalysts from a huge HEI material space (10 6 ) based on our newly established machine learning (ML) driven “decode—describe—design” (3D) framework by experimentally fabricated A 3 B‐type (FeCoNi) 3 (AlTi) system. Over 700 catalysts exhibited better performance than existing experimental results, indicating that the experiment only touched a very small part of the material space. Moreover, we developed various powerful descriptors (such as Λ OH , Λ H ) and analysis tools (such as, RDERA, RPDAD, CPMCC, CPMCS) to decouple the complex interplay of elements into atomic‐ and region‐specific effects, laying the foundation for the establishment of structure‐activity relationships and guiding the rational design of catalysts. The interpretable ML‐driven 3D framework, powerful descriptors, and novel analysis tools enable efficient design and screening, catalytic mechanism elucidation, and structure‐activity relationship establishment. They are expected to stimulate further computational and experimental investigations in related catalyst systems.

Hao Deng, Liming Yang · 0 citations
Open access Aug 2026

Decoding the chemical space of fast-ion conductors via a descriptor-guided transfer learning framework

Fast-ion conductors (FICs) are key components for next-generation high-performance batteries, yet predicting ion mobility remains challenging because of the unclear transport mechanisms. This difficulty is further compounded by experimental datasets that lack precise crystal structures. Here, we present a descriptor-guided transfer learning framework, named IonNet, to predict ion mobility for compounds, regardless of structure accessibility. IonNet adopts a multichannel subnetwork architecture that captures universal representations by integrating static and statistical descriptors of compounds. We demonstrate the exceptional performance of IonNet in predicting ion mobility, consistently outperforming 16 ablation study combinations and previous deep learning models. Leveraging the universal adaptability of chemical representations, IonNet not only uncovers 87 FICs among ∼4500 stable perfectly stoichiometric compounds but also efficiently pinpoints ∼63,000 prospective FICs from ∼5 million substituted compounds. This study not only presents a full-chain artificial intelligence tool for identifying FICs but also offers compositional principles governing ion mobility, thereby accelerating the development of energy storage and conversion.

Zhilong Wang, Fengqi You · 0 citations
Jul 2026

DFTB-Based Validation of Molecules Generated by Reversible Junction Tree Reinforcement Learning: A Case Study on Molecular Solar Thermal Fuels

Machine-learning methods are widely used for molecular generation across vast chemical spaces. However, most studies evaluate model performance using empirical descriptors such as logP and molecular weight, which do not indicate whether generated molecules satisfy target properties. Moreover, when generative models propose molecules outside existing databases—as is desirable in exploratory design—reference data are unavailable, making validation difficult. To move data-driven molecular design toward practical applications, it is essential to couple molecular generation with quantum-chemical calculations and to establish a validation workflow that screens many candidates while retaining physical reliability. Density-functional tight-binding (DFTB) balances accuracy and cost and is suitable for large-scale quantum-chemical screening. Reversible junction tree reinforcement learning (RJT-RL) generates molecules by assembling fragments on a reversible tree representation, combining interpretability with goal-directed optimization. In this work, we construct a validation framework that couples RJT-RL with DFTB calculations. Using azobenzene-based molecular solar thermal fuels (STFs) as a model system, we evaluate the property distributions of RL-generated molecules and compare them with reference molecules. On the generation side, we fix the azobenzene backbone and construct candidates by attaching substituents to the aromatic rings and, when applicable, further functionalizing them. A database containing approximately 5×10 4 azobenzene derivatives is used as an expert dataset to pretrain the RJT-RL model, allowing the policy network to learn structural patterns. In the subsequent reinforcement-learning stage, the reward is defined as the maximum Tanimoto similarity between each generated molecule and molecules in this database. Figure 1(a) shows that the maximum similarity initially fluctuates and then reaches a plateau. We therefore define the first 1–12k generated molecules as the oscillation stage and the 15–27k molecules as the platform stage, and perform DFTB calculations for all molecules generated in these two stages to compare generation behavior and properties before and after convergence of the reward. On the validation side, we apply a unified DFTB workflow to both RL-generated molecules and database molecules. SMILES strings are converted into three-dimensional trans and cis conformers using Open Babel. Geometry optimizations are then performed with DFTB+, followed by ground-state to first excited-state (S 0 →S 1 ) excitation-energy calculations on the optimized trans structures. This yields the excitation energy ΔE exc , relevant to matching the solar spectrum, and the energy difference ΔE iso between the trans and cis isomers, characterizing the energy-storage capacity. Candidates for which either geometry optimization or the excitation-energy calculation fails are excluded from the subsequent property statistics, and the overall DFTB success rate is used as a simple proxy for the structural reasonableness of the generated molecules. Successfully calculated molecules Success rate Oscillation stage (1-12k) 3100 25.8% Platform stage (15k-27k) 7992 66.6% This table summarizes the DFTB success rates in the two training stages. In the oscillation stage, DFTB calculations succeed for only 25.8% of generated molecules, whereas in the platform stage the success rate rises to 66.6%. This indicates that, under a similarity-based reward, the trained network produces a larger fraction of geometries that lie within the applicability domain of the DFTB model. Figure 1(b) compares the DFTB property distributions of reference database molecules and of molecules generated in the oscillation and platform stages. For ΔE iso , the overall distributions in the two training stages are similar to that of the database, suggesting that RL-generated molecules do not strongly deviate from the reference set in terms of energy-storage capacity. In contrast, the ΔE exc distribution of molecules from the platform stage is more concentrated than that from the oscillation stage and exhibits a pronounced peak in the energy region overlapping with the visible spectrum, indicating an enrichment of candidates in this desirable range compared with the reference database. Applying simple STF criteria, such as requiring ΔE exc to fall within a target window and ΔE iso to exceed a threshold, yields a subset of RL-generated molecules with potential relevance. At the same time, the weak correlation between the similarity-based reward and DFTB-computed ΔE exc and ΔE iso shows that the reward magnitude alone is insufficient to reliably predict property quality, highlighting the need for quantum-chemical validation of RL-generated molecules. In summary, we present a workflow that combines reversible junction tree reinforcement learning with rapid DFTB calculations, and demonstrate its use in assessing the chemical reasonableness and property distributions of similarity-driven RL-generated molecules in an azobenzene-based STF system. The framework is not limited to STF chemistry or specific target properties: by changing the reward definition and the set of properties computed with DFTB, it can be extended to other molecular design tasks, providing a general strategy for evaluating and calibrating the physical reliability of AI-generated molecules. Figure 1

Qingyu Zhu, Ryan T. Lingg, Ian McCleary et al. · 0 citations