Similar papers
Data-driven prediction and classification of multi-component alloys using interpretable machine learning
The design of advanced metallic alloys is challenged by complex, nonlinear interactions among multiple alloying elements, making conventional trial-and-error approaches costly and time-intensive. This study presents an integrated, interpretable machine learning framework applied to a dataset of 2,672 multi-component alloy systems, using elemental composition as the sole input. A multi-output Random Forest Regressor simultaneously predicts Ultimate Tensile Strength (UTS) and Liquidus Temperature, achieving a test R² of 0.848 and a 5-fold cross-validation R² of 0.849 ± 0.017, outperforming Linear Regression (R² = 0.512) and Gradient Boosting (R² = 0.787) baselines. A Logistic Regression classifier identifies high-performance alloy compositions defined by simultaneous UTS and liquidus temperature thresholds achieving an overall accuracy of 82% and an AUC of 0.871. Threshold sensitivity analysis confirms classification stability across varying performance criteria. Principal Component Analysis (PCA) and K-Means clustering, applied to the full compositional feature space, reveal three distinct alloy families with systematic differences in mechanical and thermal properties. Critically, interpretability is preserved throughout via SHAP analysis and permutation-based feature importance, identifying Vanadium (V), Iron (Fe), Tungsten (W), and Carbon (C) as the dominant compositional drivers a finding that both confirms and extends established metallurgical understanding. The dataset was verified to contain no missing values across all 31 elemental features. Collectively, the proposed framework provides a scalable and interpretable foundation for accelerated alloy screening, property prediction, and data-driven materials discovery.
Flowability prediction and artificial intelligence analysis of SCC based on machine learning
Six machine learning algorithms were employed to construct artificial intelligence models for the precise prediction of self-compacting concrete (SCC) flow properties, and the extreme gradient boosting (XGB) model was identified as exhibiting superior predictive accuracy and generalization performance.
Validation of a Machine Learning–Assisted LIBS Model for Quantitative Steel Analysis
Laser-induced breakdown spectroscopy (LIBS) shows great promise for the rapid chemical characterization of materials. However, quantitative analysis of elements remains challenging due to strong matrix effects and the predominance of emission lines. In the present study, a machine learning–assisted LIBS model (LIBS-ML pipeline) was employed to analyze carbon, medium-alloy, and high-alloy steels. A total of 900 spectra from 18 reference specimens were used to train Random Forest (RF), Gradient Boosting (GB), and Extremely Randomized Trees (ET) ensemble models, while five independent steel specimens were reserved for prediction evaluation. The ET model demonstrated the best training performance (MSE = 0.1551; R² = 0.9435), while the RF model exhibited greater stability during independent validation. Low prediction errors were obtained for carbon steels, with mean absolute error (MAE) values as low as 0.0142 wt% for C and 0.0178 wt% for Mn. In medium-alloy steel, the predicted values of Cr and Ni remained close to the nominal compositions. Higher deviations were observed in high-alloy steel, with MAE reaching 1.1191 wt% for Ni and 1.0919 wt% for Mo, reflecting the increased complexity of the matrix. The obtained results confirmed the LIBS–ML pipeline applicability for quantitative steel analysis and highlight its potential as a rapid tool for metallurgical monitoring and alloy characterization.
Phase-stability-guided explainable machine learning for mechanical-property prediction and experimentally supported candidate screening of high-entropy alloys
High-entropy alloys (HEAs) provide a broad compositional space for developing structural materials with balanced phase stability and mechanical performance. However, reliable mechanical-property prediction remains challenging because alloy chemistry, phase constitution, processing state, and model uncertainty are strongly coupled. We developed a phase-stability-guided explainable machine learning framework using a curated database comprising 541 phase-labelled HEAs, 263 hardness-labelled records, and 214 yield-strength-labelled records. Hierarchical physical and processing descriptors were combined with leakage-controlled phase-probability features generated through composition-level nested cross-fitting. XGBoost and CatBoost models were used for property prediction, bootstrap ensembles for uncertainty quantification, and SHAP and accumulated local effects for model interpretation. A total of 300,000 virtual candidates were screened, followed by CALPHAD-assisted assessment and experimental validation of three representative alloys. The phase classifier achieved an overall accuracy of 0.923. The phase-guided property models achieved R² values of 0.881 for hardness and 0.856 for yield strength. The experimentally measured dominant phases, hardness values, and compressive yield strengths of three representative candidates were generally consistent with the model predictions, with moderate experimental deviations. The proposed framework integrates physical descriptors, probabilistic phase information, model explainability, and uncertainty-aware screening, providing an interpretable and experimentally supported strategy for prioritizing promising HEA compositions before broader experimental optimization.
Interpretable Deep Learning–Machine Learning Models for Accelerated Discovery of Elastic Properties in Refractory High‐Entropy Alloys
Refractory high‐entropy alloys (RHEAs) offer mechanical performance at extreme temperatures, but their optimization is hindered by the vast compositional space and high cost of first‐principles calculations. This study establishes an interpretable machine‐learning benchmark for predicting nine elastic and mechanical properties of RHEAs. An EMTO‐CPA dataset comprising 2487 alloys is used to evaluate eight regression models: Gaussian process regression (GPR), support vector regression (SVR), shallow and deep neural networks, LightGBM, XGBoost, CatBoost, and Histogram‐based Gradient Boosting. Seven composition‐derived descriptors representing atomic size, electronic structure, and thermodynamic characteristics are employed. Hyperparameters are optimized using randomized search with five‐fold cross‐validation, while robustness is assessed through ten train–test evaluations using random seeds 42–51. Model performance is reported as mean ± standard deviation of MAE and R 2 . The results demonstrate target‐dependent performance: GPR performs particularly well for sws, SVR achieves favorable performance for several properties, and deep neural networks provide strong predictions for selected mechanical targets. SHAP analysis of optimized SVR models identifies valence electron concentration, average atomic radius, atomic size mismatch, and melting temperature as important contributors, whereas mixing entropy and mixing enthalpy generally show weaker contributions. Overall, the benchmark provides a reproducible framework for RHEA prediction and supports screening and alloy design.
Interpretable Machine Learning for Mechanical Property Prediction of 5Cr-0.5Mo Steel: SHAP Explainability, Multi-Model Comparison, and Uncertainty Quantification
5Cr-0.5Mo ferritic steels are widely used in high-temperature power-plant components. Although artificial neural network (ANN) models have shown good performance in predicting tensile properties, they provide limited insight into predictions and generally do not quantify the uncertainty. In this study, three tree-based machine learning models—Random Forest (RF), XGBoost (XGB), and Gradient Boosting (GB)—were developed using 36 unique alloy grade–temperature observations from a validated NIMS 5Cr-0.5Mo tensile dataset. The model performance was evaluated using leave-one-grade-out (LOGO) cross-validation, with pooled out-of-fold (OOF) predictions used to assess the overall performance. SHapley Additive exPlanations (SHAP) were used to examine feature contributions, whereas Gaussian Process Regression (GPR) was evaluated as a proof-of-concept for uncertainty quantification of yield strength (YS). GB showed the strongest performance for ultimate tensile strength (UTS) and reduction in area (RA), achieving pooled OOF R2 values of 0.9698 and 0.9570, respectively. RF achieved corresponding R2 values of 0.9406 and 0.9488, respectively. SHAP identified the test temperature as the most influential feature across all four properties, whereas the Cr content and austenite grain size contributed significantly to the strength predictions. For YS, the GPR achieved complete empirical coverage of the 95% predictive intervals, although the relatively large mean interval width indicated conservative uncertainty estimates. Given the limited dataset and feature correlations, the SHAP results should be regarded as exploratory, rather than mechanistic. Overall, this study demonstrates the potential of interpretable, uncertainty-aware ML for small alloy datasets, while emphasizing the need for larger, compositionally diverse datasets and independent validation.