Skip to content

Similar papers

Open access Jul 2026

Prediction of solubility of hydrogen in chemicals using QSPR-based machine learning approach: a comparative study

The ‘Quantitative Structure–Property Relationship’ (QSPR) method has been used for the prediction of solubility of hydrogen (x) in different chemicals. The dataset consists of 3761 datapoints including 100 unique chemicals at the wide ranges of T and P. An MLR-model, the simplest form of machine learning algorithm, was constructed using the selected descriptors to predict x in various chemicals. For the first time, a comprehensive and predictive MLR-based QSPR model has been developed for this target. The dataset was divided into a training set including 2570 datapoints and to a test set including 1191 datapoints. The advantages of this approach are thoroughly discussed and compared with other available models which were developed with other ML algorithms. Unlike previous models, internal validation was performed on the MLR-QSPR model. According to the results of statistical parameters (R2 = 0.96 and Q2LOO-CV = 0.96), the predictive capability of the MLR-QSPR model was acceptable for training set.

Ali Ebrahimpoor Gorji, V. Alopaeus · 0 citations
Aug 2026

Machine learning prediction of organic compound melting points informed by condensed-phase and electronic descriptors.

Melting point (MP) is an important thermophysical property for the chemical process industry, yet accurate prediction of MP for organic compounds in the absence of experimental data remains challenging due to the complex interplay between molecular packing, intermolecular interactions, and electronic structure. Traditional group contribution and quantitative structure-property relationship models, which rely primarily on static molecular descriptors, often fail to capture these critical condensed-phase effects. In this study, we present a hybrid machine learning framework that integrates cheminformatics descriptors with quantum chemical features and dynamic condensed-phase descriptors derived from molecular dynamics (MD) simulations. Using a curated subset of the DIPPR 801 database, multiple machine learning architectures, including light gradient boosting machine (LightGBM) and graph convolutional networks, were evaluated with feature sets of increasing physical fidelity. The best-performing model, based on LightGBM trained on Dragon descriptors augmented with MD and quantum chemical features, achieves a mean absolute error of 22.5 K, outperforming descriptor-only models and structure-based deep learning baselines. Shapley additive explanations interpretability analysis reveals that melting behavior is governed primarily by molecular topology, surface-area-weighted electronic descriptors, and condensed-phase interaction properties. In contrast, many isolated functional group and single molecule electronic descriptors contribute negligibly once these effects are accounted for. These results demonstrate that incorporating physics-informed, multi-scale descriptors enables more accurate and physically interpretable MP predictions.

Frank T. Mtetwa, N. Giles, W. Wilding et al. · 0 citations
Open access Jul 2026

Machine learning-guided screening and validation of antioxidant small molecules from Clausena lansium

To discover natural antioxidants from Clausena lansium (Lour.) Skeels for food applications, we established a 474-compound database and applied a multiscale workflow integrating ensemble machine learning, molecular docking, structural clustering, molecular dynamics simulations, and quantum chemical calculations. The ensemble models achieved ROC-AUC values of 0.85–0.99 across eight antioxidant assays. Multi-criteria screening yielded 47 candidates, and 30 structurally diverse candidates underwent molecular dynamics simulations. Quercetin 3-arabinoside was prioritized, exhibiting stable non-covalent binding within the Keap1 Kelch domain during a 500 ns simulation. Quantum chemical calculations showed a HOMO–LUMO gap of 4.0589 eV, 25.5% lower than that of vitamin C. Given standard availability, its structural isomer Avicularin was assessed and exhibited dose-dependent antioxidant activity, with DPPH and ABTS IC50 values of 26.84 and 6.94 μM, respectively, and a FRAP value of 1.327 ± 0.218 mmol FeSO4 equivalents L−1 at 25 μM. This study supports antioxidant discovery from edible fruits.

Yufan Tong, Min Wang · 0 citations
Aug 2026

Detour Ring Index as a Molecular Descriptor for Predicting Gas Chromatographic Retention of Sulfur Polycyclic Aromatics

Prediction of positional isomers in polycyclic aromatic sulfur compounds remains a challenge in analytical chemistry. Here, simple ring‐based topological descriptors derived from Distance/Detour Ring Indices (D/DIr05, D/DIr06, and D/DIr09) are evaluated for benzothiophene and dibenzothiophene positional isomers. D/DIr06 and D/DIr09 show excellent linear correlations with experimental gas chromatographic retention indices ( R 2 = 0.992 and 0.991, respectively), while D/DIr05 exhibits a weak correlation ( R 2 = 0.325), suggesting its utility for classification rather than prediction. The strong performance of D/DIr06 and D/DIr09 reflects their sensitivity to structural variations in ring fusion and substitution patterns. These descriptors offer a simple, efficient, and cost‐effective means for predicting retention behavior and identifying positional isomers, providing valuable tools for quantitative structure–retention relationship studies of sulfur‐containing aromatic compounds.

N. E. Moustafa, Kout El‐Kloub Fars Mahmoud, Abdulelah H. Alsulami et al. · 0 citations
Open access Jul 2026

Performance Prediction of Metal Nitride Energetic Materials Based on Structural Searching and Machine Learning

Metal nitrides exhibit broad prospects as high-energy-density materials (HEDMs) due to their exceptional energy release characteristics and environmentally friendly decomposition products. However, a fundamental challenge for this class of HEDMs lies in the trade-off between high energy density and high stability. Here, we employed a crystal structure search method to obtain numerous metal nitride configurations with different elements and stoichiometries and characterized their properties using first-principles calculations and ab initio molecular dynamics (AIMD) simulations. Based on the dataset, we developed descriptors related to the elemental and structural properties and trained regression models to predict the energy density of the metal nitrides. Additionally, multiple classifier models were trained to assess their stability. Through these machine learning models, we analyzed the features affecting the energy density and stability of metal nitrides and identified three key descriptors: the average distance between each nitrogen atom and its nearest neighbor (DNN), the stoichiometric ratio (N:M), and the average number of N-N bonds per nitrogen atom (aveNNBonds). Based on these insights, we propose a design principle for advanced metal nitride HEDMs: prioritizing high nitrogen-to-metal ratios, light metal elements, and structures wherein nitrogen atoms are spatially separated by the metal matrix, which minimizes N-N bonds and favors dominant M-N bonding. Moreover, our analysis revealed a class of metal nitrides with high nitrogen content that combines considerable stability and high energy density, demonstrating the power of machine learning in material HEDM design.

Yaozhong Liu, Huifang Du, Caimu Wang et al. · 0 citations
Open access Aug 2026

Prediction of the Molecular Lipophilicity of an Alkylphenol Family Using Quantum Chemistry and QSPR Methods

This study examined the relationship between the water/octanol partition coefficient of eighteen alkylphenols and molecular descriptors derived from quantum-chemical calculations using a quantitative structure-property relationship (QSPR) approach. The experimental database was divided into a training set of fourteen compounds and a test set of four compounds. Three descriptors, the electronic energy (ET), the energy of the Lowest Unoccupied Molecular Orbital (ELUMO) and the energy gap (ΔEgap) were used to develop a multiple linear regression model. The model showed a strong association between the experimental lipophilicity values and the selected descriptors, with R = 0.9850, R² = 0.9703, a standard deviation of 0.0853, and F = 108.9532. Internal validation using leave-one-out cross-validation and property randomisation, together with external validation using the test set and the Tropsha criteria, was applied to assess model stability and predictive performance. The applicability domain was evaluated using a Williams plot. Within the studied dataset and the defined applicability domain, the model showed close agreement between experimental and predicted water/octanol partition coefficients. The reported validation results indicate that the selected quantum-chemical descriptors can be used to model lipophilicity within this alkylphenol series. Predictions for additional alkylphenols should, however, remain restricted to compounds that fall within the model’s defined applicability domain.

Fatogoma Diarrassouba, K. Bamba, N. Ziao · 0 citations