Skip to content

Similar papers

Jul 2026

Predicting soil-water partition coefficients of PFAS using machine learning: Model development, interpretation, and validation.

PER: and polyfluoroalkyl substances (PFAS) are environmentally persistent contaminants, yet experimental determination of their soil-water partition coefficients (Kd) remains costly and time-consuming. In this study, five machine-learning regression models were developed using 2057 literature-derived batch adsorption data points by integrating the average net charge descriptor (Zavg), the composite descriptor alert_prior_score, equilibrium aqueous concentration (Cw), PFAS structural descriptors, and soil physicochemical properties. Among the tested models, extreme gradient boosting (XGBoost) showed the best performance with 10 descriptors, achieving R2 values of 0.83 and 0.86 and ratio of performance to deviation (RPD) values of 2.42 and 2.66 for the Cw < 10 μg/L and Cw ≥ 10 μg/L datasets, respectively. Shapley additive explanations (SHAP) analysis indicated that hydrophobic interactions dominated adsorption at low concentrations (Cw < 10 μg/L), whereas headgroup-related hydrophilicity became more influential at higher concentrations. Independent sorption experiment validation using contaminated site soils showed that prediction deviations for all samples were within one order of magnitude. These results demonstrate that the proposed model provides an efficient and interpretable tool for predicting PFAS soil-water partitioning and supports environmental risk assessment and contaminated-site management.

Yue Zhou, Hao Chen, Xi Wang et al. · 0 citations
Review Aug 2026

Quantitative structure-retention relationships (QSRR): Effect of experimental variables on QSRR model performance in HPLC - A critical review.

The QSRR modeling framework enables researchers to use molecular descriptors for predicting chromatographic retention, including HPLC retention based on their physicochemical characteristics, as it can predict chromatographic retentions and not only HPLC retention. QSRR models show low transferability between different laboratories and instruments and experimental protocols because their performance depends on experimental conditions which remain poorly documented. The present study investigates the way experimental conditions affect both model architecture and descriptor selection through their experimental design which has not been studied in previous reviews. The review demonstrates how organic modifier type and concentration, mobile phase pH and buffer identity, stationary phase chemistry, and column temperature functions as the fundamental drivers of analyte retention which leads to QSRR model transferability problems. The research examines how dataset variation and laboratory differences and variable relationships affect study results. The study evaluates traditional modeling methods which include Linear Solvation Energy Relationships and Multiple Linear Regression and Partial Least Squares together with modern machine learning techniques which use Random Forests and Gradient Boosting and Support Vector Regression and Graph Neural Networks. The research establishes model validation standards which include applicability domain assessment and descriptor selection and standardized reporting.

Mayur Borase, Kunal Patil, P. Ghode · 1 citation
Open access Jul 2026

Prediction of solubility of hydrogen in chemicals using QSPR-based machine learning approach: a comparative study

The ‘Quantitative Structure–Property Relationship’ (QSPR) method has been used for the prediction of solubility of hydrogen (x) in different chemicals. The dataset consists of 3761 datapoints including 100 unique chemicals at the wide ranges of T and P. An MLR-model, the simplest form of machine learning algorithm, was constructed using the selected descriptors to predict x in various chemicals. For the first time, a comprehensive and predictive MLR-based QSPR model has been developed for this target. The dataset was divided into a training set including 2570 datapoints and to a test set including 1191 datapoints. The advantages of this approach are thoroughly discussed and compared with other available models which were developed with other ML algorithms. Unlike previous models, internal validation was performed on the MLR-QSPR model. According to the results of statistical parameters (R2 = 0.96 and Q2LOO-CV = 0.96), the predictive capability of the MLR-QSPR model was acceptable for training set.

Ali Ebrahimpoor Gorji, V. Alopaeus · 0 citations
Jul 2026

Filling Ecotoxicity Data Gaps for Plasticizers via an Integrated Framework: Toxicity Prediction, Prioritization, and Hazardous Concentration Derivation.

Plasticizers are widely used polymer additives whose release into aquatic environments poses ecological risks, while their structural diversity challenges conventional toxicological assessment. To address this data gap, an integrated modeling framework was established for hazard assessment and prioritization of plasticizers. Species-specific QSAR models were developed for six species across three trophic levels, enabling accurate toxicity prediction for 600 plasticizers. Eighty-five plasticizers were identified as high-priority due to potent toxicity (LC50 < 10 μM) across trophic levels, 54.1% of which are alternative plasticizers. Approximately half of these high-priority plasticizers are widely used in polyvinyl chloride plastics, and 21.2% occur in bioplastics, indicating their potential contribution to the chemical risk of conventional plastics and "green" polymers. Across multiple taxa, predicted toxicity differences among plasticizer subclasses were linked to structural features captured by 2D descriptors. By integrating QSAR predictions with interspecies correlation estimation, the hazardous concentrations for 5% of species (HC5: 0.04-274 μg/L) were derived for high-priority plasticizers using species sensitivity distributions, with crustaceans identified as the most sensitive taxa. Linking hazard prediction and prioritization with polymer occurrence enables the systematic identification of high-priority plasticizers and highlights polymers that contain plasticizers of potential concern.

Jingya Li, Weigang Liang, Lin Niu et al. · 0 citations
Open access Aug 2026

Small-Sample Prediction and Uncertainty Assessment of Soil Organic Carbon Content in Cropland of the Liaohe Plain Based on the TabPFN Model

Soil organic carbon (SOC) is a key indicator of cropland quality, soil fertility, and the carbon sequestration potential of agroecosystems. Accurate characterization of its spatial distribution is essential for black soil conservation and regional soil carbon management. However, regional-scale SOC prediction is often constrained by limited field observations, which can reduce model generalizability and predictive reliability. In this study, we developed a limited-sample SOC prediction framework for the Liaohe Plain using 310 surface (0–20 cm) soil samples collected in 2025 and multi-source environmental covariates, including climate, vegetation, soil spectral, and terrain variables. The framework used the Tabular Prior-Data Fitted Network (TabPFN), whose performance was compared with that of Random Forest, Support Vector Machine, CatBoost, K-Nearest Neighbors, and XGBoost. Model performance was evaluated using 100 repetitions of random 80:20 holdout validation and repeated five-fold spatial cross-validation based on spatially constrained clustering, while sampling-induced relative uncertainty was quantified using 100 repeated random sampling and model-fitting runs. Under random holdout validation, TabPFN showed competitive predictive performance, with mean R2 and RMSE values of 0.608 ± 0.038 and 4.026 ± 0.184 g kg−1, respectively. Repeated spatial cross-validation yielded more conservative performance estimates, with mean R2 and RMSE values of 0.535 ± 0.072 and 4.33 ± 0.31 g kg−1, respectively, indicating that random splitting may overestimate model performance when sampling sites are spatially clustered. Spatial prediction showed that cropland SOC ranged from 4.63 to 27.04 g kg−1, with generally lower values in the west and higher values in the northeast. Areas with high sampling-induced relative uncertainty were mainly concentrated in the northern, northeastern, and marginal regions. These findings provide a methodological basis for SOC mapping, supplementary sampling optimization, and regional soil carbon management under limited-sample conditions, although the temporal robustness of the results requires confirmation using independent data from additional years.

Yongqiang Yang, Yanzhi Zhao, Shuang Gang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.