Aug 2026· Journal of Environmental Management· Vol 415, pp.
130684
· 0 citations· 38 references
Medicine
Abstract
Accurate prediction of second-order rate constants (k) for reactions between contaminants and reactive species (RS) is essential for understanding transformation pathways and optimizing advanced oxidation/reduction processes (AOPs/ARPs). However, existing QSAR models mostly rely on manually engineered molecular descriptors and have limited capability in capturing complicated structure-reactivity relationships, while the application of pre-trained Transformer-based molecular language models in k prediction remains largely unexplored. In this study, a deep learning-based QSAR framework (ChemBERTa-FC) was developed to predict k values for reactions of water contaminants with HO•, SO4•-, and eaq-. The model integrates a pre-trained molecular language model (ChemBERTa) for representation learning with Fully Connected (FC) layers for regression, enabling end-to-end prediction directly from SMILES without manual feature engineering. The proposed model achieves high predictive performance across all three RS: for HO•, the model yields RMSE values of 0.052 (training) and 0.088 (test); for SO4•-, RMSEtest and R2test reach 0.108 and 0.586, respectively; for eaq-, the model exhibits consistently low error and balanced performance across the full reactivity range. SHAP analysis reveals RS-specific attribution patterns aligned with reaction mechanisms, while Pearson correlation shows that most learned embeddings are not linearly explainable by traditional descriptors, indicating the capture of higher-order structure-activity relationships. Applicability domain analysis further confirms the reliability of model predictions within the defined chemical space. Overall, this work establishes a transferable and interpretable deep learning framework for k prediction and provides new insights into the molecular determinants of contaminant reactivity, supporting the rational design of water treatment processes.
Selective conversion of nitrogen-containing species into harmless molecular nitrogen (N2) remains a key challenge in the catalytic oxidation of nitrogen-containing volatile organic compounds (NVOCs). Machine learning (ML) provides an effective approach for predicting catalytic performance and identifying key descriptors from complex literature-derived datasets. Herein, a literature-derived catalyst database was constructed to predict N2 selectivity during NVOC oxidation and clarify the factors governing nitrogen transformation. Thirteen descriptors related to catalyst composition, structural properties, support acidity, and reaction conditions were used to train eight ML models. Among them, the ExtraTrees model exhibited the best predictive performance, with a coefficient of determination of 0.958 and a root mean square error of 7.638 on the test set. Shapley additive explanations and partial dependence plots revealed that oxygen concentration, reactant concentration, reaction temperature, gas hourly space velocity, and support acidity were the dominant factors affecting N2 selectivity, with support acidity identified as the key catalyst-related descriptor. Guided by this descriptor-level insight, Cu/M and CuFe/M catalysts (M = SiO2, ZSM-5, and Al2O3) were prepared and evaluated for acetonitrile oxidation. The catalytic and spectroscopic results confirmed the predicted role of support acidity, showing that different supports regulate CH3CN adsorption, CN activation, and the evolution of hydrolysis and oxidation related nitrogen-containing intermediates, thereby affecting nitrogen-product distributions and N2 selectivity. This work integrates interpretable machine-learning prediction with targeted external validation and mechanistic analysis, providing mechanistic insight into support-acidity-regulated nitrogen transformation and guidance for designing NVOC oxidation catalysts with high N2 selectivity.
Haotian Hu, Ying Wang, Zihao Zhai et al.· Journal of Colloid and Inter...· 0 citations
The ‘Quantitative Structure–Property Relationship’ (QSPR) method has been used for the prediction of solubility of hydrogen (x) in different chemicals. The dataset consists of 3761 datapoints including 100 unique chemicals at the wide ranges of T and P. An MLR-model, the simplest form of machine learning algorithm, was constructed using the selected descriptors to predict x in various chemicals. For the first time, a comprehensive and predictive MLR-based QSPR model has been developed for this target. The dataset was divided into a training set including 2570 datapoints and to a test set including 1191 datapoints. The advantages of this approach are thoroughly discussed and compared with other available models which were developed with other ML algorithms. Unlike previous models, internal validation was performed on the MLR-QSPR model. According to the results of statistical parameters (R2 = 0.96 and Q2LOO-CV = 0.96), the predictive capability of the MLR-QSPR model was acceptable for training set.
Ali Ebrahimpoor Gorji, V. Alopaeus· Journal of Computer-Aided Mo...· 0 citations
Experimental determination of flash points (FPs) for liquid mixtures is laborious and costly, highlighting the need for reliable predictive approaches for safety assessment and engineering applications. Although numerous models have been reported for binary miscible mixtures, most rely on fixed model parameters or empirical correlations, which limits their ability to capture the nonlinear relationship between molecular structure and FP. In this study, a quantitative structure-property relationship (QSPR) framework that tightly integrates differential evolution (DE) with support vector regression (SVR) was developed to predict the FP values of binary miscible mixtures, where DE was employed to globally optimize key SVR hyperparameters and enhance model generalization capability. A dataset consisting of 332 compositions from 33 binary mixtures formed by pairwise combinations of 20 pure components was employed, and multiple molecular descriptor representation strategies were deliberately adopted to construct distinct DE-SVR models, enabling a systematic investigation of the combined effects of descriptor representation and model optimization on predictive performance. Three DE-SVR models were established based on different descriptor sets, and their predictive accuracy, robustness, and stability were comprehensively evaluated. Among them, the model constructed using physicochemical parameters exhibited the best overall performance. Comparative analyses with existing FP prediction methods reported in the literature further confirmed the effectiveness and superiority of the proposed DE-SVR-based models. The results of this study provide a practical tool for FP estimation of binary mixtures and offer valuable insights into the joint roles of model optimization and molecular representation in mixture property prediction.
Shuangyu Song, Xiaoya Song, Xingqian Chen et al.· Journal of Molecular Graphic...· 0 citations
Melting point (MP) is an important thermophysical property for the chemical process industry, yet accurate prediction of MP for organic compounds in the absence of experimental data remains challenging due to the complex interplay between molecular packing, intermolecular interactions, and electronic structure. Traditional group contribution and quantitative structure-property relationship models, which rely primarily on static molecular descriptors, often fail to capture these critical condensed-phase effects. In this study, we present a hybrid machine learning framework that integrates cheminformatics descriptors with quantum chemical features and dynamic condensed-phase descriptors derived from molecular dynamics (MD) simulations. Using a curated subset of the DIPPR 801 database, multiple machine learning architectures, including light gradient boosting machine (LightGBM) and graph convolutional networks, were evaluated with feature sets of increasing physical fidelity. The best-performing model, based on LightGBM trained on Dragon descriptors augmented with MD and quantum chemical features, achieves a mean absolute error of 22.5 K, outperforming descriptor-only models and structure-based deep learning baselines. Shapley additive explanations interpretability analysis reveals that melting behavior is governed primarily by molecular topology, surface-area-weighted electronic descriptors, and condensed-phase interaction properties. In contrast, many isolated functional group and single molecule electronic descriptors contribute negligibly once these effects are accounted for. These results demonstrate that incorporating physics-informed, multi-scale descriptors enables more accurate and physically interpretable MP predictions.
Frank T. Mtetwa, N. Giles, W. Wilding et al.· Journal of Chemical Physics· 0 citations