Quantitative Prediction of the Transformation Potential of Polyfluoroalkyl Substances: A Computational and Machine Learning Study of •OH‑Initiated Initial Reaction Kinetics
Jun 2026· Communications in Computational Chemistry· Vol 8, pp. 300-307· 0 citations
TL;DR
This work establishes a robust predictive framework for the •OH-initiated initial transformation potential of polyfluoroalkyl substances, providing a high-throughput tool for the environmental risk assessment and preliminary screening of polyfluoroalkyl alternatives with controlled transformation behavior.
Abstract
Per- and polyfluoroalkyl substances (PFAS) are ubiquitous persistent organic pollutants, and hydroxyl radical (•OH)-initiated abiotic transformation dominates the environmental fate of polyfluoroalkyl substances. However, quantitative relationships between molecular structures and the •OH-initiated transformation reactivity remain poorly understood. Herein, we combined density functional theory (DFT) calculations and machine learning (ML) modeling to establish a predictive framework for the •OH-initiated initial transformation potential of polyfluoroalkyl substances. The Gibbs free energy barriers of the rate-determining hydrogen atom abstraction (HAA) step were computed at the M06-2X/6-311++G(2d,2p) level with the SMD solvation model for polyfluoroalkyl substances with varied functional groups, perfluoroalkyl chain lengths, C-H bond positions, alkyl spacers, and branched configurations. The results revealed that perfluoroalkyl chain length imposed negligible effects on HAA barriers, whereas functional groups governed the transformation reactivity in the order of $-P{O_4}^{2-}>-OH≈-CO{O}^->-S{O_3}^-.$ The positive charge density of aliphatic C-H hydrogens served as a decisive descriptor, in which stronger C-H polarization induced higher HAA barriers. Alkyl spacers shielded the electron-withdrawing effect of fluorine, while branched CF moieties enhanced C-H polarization. The optimal machine learning model (RF-MACCS) achieved excellent predictive performance $(R^2=0.91),$ and its practical application on environmentally relevant polyfluoroalkyl substances yielded predictions highly consistent with DFT results. SHAP analysis verified the mechanistic consistency between model predictions and DFT-derived structure-reactivity relationships. This work establishes a robust predictive framework, providing a high-throughput tool for the environmental risk assessment and preliminary screening of polyfluoroalkyl alternatives with controlled transformation behavior.
Proton-coupled electron transfer (PCET) mediated by hydroquinone and related molecules is key to natural and artificial energy conversion. The reactivity of these molecules depends on their bond dissociation free energy (BDFE), but studying the relationship between structure and thermochemistry across this chemical space has been limited by challenging experimental setup and high computational expense. Here, we present the first use of the AIMNet2 neural network potential to calculate average BDFE (BDFEavg) values for the 2H+/2e− dehydrogenation of about 200 000 hydroquinone-like compounds, including vicinal diamines, diols, and dithiols. Benchmarking against DFT calculations for 168 substituted ortho-phenylenediamines (opda) shows good agreement (R2 ∼ 0.84). Our analysis finds that the BDFEavg of diamines ranges from 50 to 80 kcal mol−1 and can be systematically tuned by modifying the backbone and N-substitution: electron-withdrawing groups raise BDFEavg by up to 15 kcal mol−1, while lower aromaticity in furan and thiophene backbones decreases BDFEavg by approximately 10 kcal mol−1 compared to the phenyl systems (∼65 kcal mol−1). Validation through cyclic voltammetry and reactivity studies with quinone oxidants for selected compounds supports the computational results. This extensive thermochemical database and a web-based prediction tool developed as a result of this work will offer valuable resources for designing PCET reagents for catalysis, energy storage, and biomedical uses.
Rajdeep Sarma, Yiwen Wang, David D Hebert et al.· Chemical Science· 0 citations
Fluorinated molecules, including per- and polyfluoroalkyl substances (PFAS), present persistent challenges for thermochemical characterization due to limited experimental data, strong carbon–fluorine bonding, and the rapidly expanding size and diversity of fluorinated chemical space. While density functional theory (DFT) calculations can provide useful thermochemical data for individual fluorinated species, their routine application becomes increasingly impractical as molecular size, conformational complexity, and the number of distinct PFAS compounds continue to grow. Existing Benson-type group additivity schemes provide limited resolution for fluorinated environments, restricting their applicability to modern fluorinated and PFAS-relevant systems. Here, we develop a chemically resolved group additivity (GA) framework for fluorinated and PFAS-relevant species by fragmenting DFT-derived thermochemistry for 3070 molecules. This approach expands the available fluorinated Benson-type group library from 14 to 159 local environments and integrates the resulting groups within the Python Group Additivity (pGrAdd) framework. 10-fold cross-validated regression against DFT data yields root-mean-square deviations (RMSDs) of 8.14 kcal·mol–1 for enthalpy and 10.05 cal·mol–1·K–1 for entropy, which are reduced to 2.55 kcal·mol–1 and 6.36 cal·mol–1·K–1, respectively, following application of independently defined nongroup interaction correction terms in pGrAdd. Comparison with available experimental thermochemical data shows improved agreement and reduced bias compared to legacy Benson group libraries. This expanded fluorinated GA framework enables scalable and chemically interpretable thermochemical predictions for fluorinated and PFAS-relevant species, supporting kinetic modeling and mechanistic studies where direct electronic structure calculations are feasible but not scalable.
Samuel Eccles, Steven Pellizzeri· Journal of Physical Chemistr...· 0 citations
Accurate prediction of second-order rate constants (k) for reactions between contaminants and reactive species (RS) is essential for understanding transformation pathways and optimizing advanced oxidation/reduction processes (AOPs/ARPs). However, existing QSAR models mostly rely on manually engineered molecular descriptors and have limited capability in capturing complicated structure-reactivity relationships, while the application of pre-trained Transformer-based molecular language models in k prediction remains largely unexplored. In this study, a deep learning-based QSAR framework (ChemBERTa-FC) was developed to predict k values for reactions of water contaminants with HO•, SO4•-, and eaq-. The model integrates a pre-trained molecular language model (ChemBERTa) for representation learning with Fully Connected (FC) layers for regression, enabling end-to-end prediction directly from SMILES without manual feature engineering. The proposed model achieves high predictive performance across all three RS: for HO•, the model yields RMSE values of 0.052 (training) and 0.088 (test); for SO4•-, RMSEtest and R2test reach 0.108 and 0.586, respectively; for eaq-, the model exhibits consistently low error and balanced performance across the full reactivity range. SHAP analysis reveals RS-specific attribution patterns aligned with reaction mechanisms, while Pearson correlation shows that most learned embeddings are not linearly explainable by traditional descriptors, indicating the capture of higher-order structure-activity relationships. Applicability domain analysis further confirms the reliability of model predictions within the defined chemical space. Overall, this work establishes a transferable and interpretable deep learning framework for k prediction and provides new insights into the molecular determinants of contaminant reactivity, supporting the rational design of water treatment processes.
M. Tan, Zhouji Wu, Feng Wu et al.· Journal of Environmental Man...· 0 citations
Machine learning potential-driven molecular dynamics simulations (ML-MD) were employed to provide atomistic insights into the dehydrogenation kinetics of pristine and doped MgH2. Through systematic investigation of distinct surface orientations, the MgH2 (100) surface was identified as the most active low-index surface for hydrogen release. For pristine MgH2, our simulations revealed a novel H2 formation mechanism characterized by H2 generation in the subsurface region followed by diffusion to the surface for desorption, highlighting the critical role of subsurface processes beyond conventional surface-driven pathways. Comprehensive screening of 22 doping elements identified Ni as the most effective dopant. Among several descriptors, machine learning analysis identified the time-coupled Miedema electron density as the critical descriptor, underscoring the role of electronic properties. Consequently, a volcano-shaped relationship was uncovered between the intrinsic Miedema electron density ( nws ) and total hydrogen release (optimal window: 4.0<nws<5.4x10-2 e/bohr3). Dopants within this range serve a dual function: acting as thermodynamic sinks for H attraction while maintaining a balanced interaction strength to facilitate H-H coupling and H2 release. This atomistic-level validation provides strong theoretical support for the experimentally observed"hydrogen pump"effect of catalytic phases. The present study demonstrates the strong capability of ML- MD in navigating through complex catalytic mechanisms and establishing quantitative property-activity relationships, providing a robust framework for rational design of high-performance catalysts for MgH2 and other hydrogen storage materials.
Bo Han, Jianchuan Wang, Rui Zhang et al.· 0 citations
Selective conversion of nitrogen-containing species into harmless molecular nitrogen (N2) remains a key challenge in the catalytic oxidation of nitrogen-containing volatile organic compounds (NVOCs). Machine learning (ML) provides an effective approach for predicting catalytic performance and identifying key descriptors from complex literature-derived datasets. Herein, a literature-derived catalyst database was constructed to predict N2 selectivity during NVOC oxidation and clarify the factors governing nitrogen transformation. Thirteen descriptors related to catalyst composition, structural properties, support acidity, and reaction conditions were used to train eight ML models. Among them, the ExtraTrees model exhibited the best predictive performance, with a coefficient of determination of 0.958 and a root mean square error of 7.638 on the test set. Shapley additive explanations and partial dependence plots revealed that oxygen concentration, reactant concentration, reaction temperature, gas hourly space velocity, and support acidity were the dominant factors affecting N2 selectivity, with support acidity identified as the key catalyst-related descriptor. Guided by this descriptor-level insight, Cu/M and CuFe/M catalysts (M = SiO2, ZSM-5, and Al2O3) were prepared and evaluated for acetonitrile oxidation. The catalytic and spectroscopic results confirmed the predicted role of support acidity, showing that different supports regulate CH3CN adsorption, CN activation, and the evolution of hydrolysis and oxidation related nitrogen-containing intermediates, thereby affecting nitrogen-product distributions and N2 selectivity. This work integrates interpretable machine-learning prediction with targeted external validation and mechanistic analysis, providing mechanistic insight into support-acidity-regulated nitrogen transformation and guidance for designing NVOC oxidation catalysts with high N2 selectivity.
Haotian Hu, Ying Wang, Zihao Zhai et al.· Journal of Colloid and Inter...· 0 citations