This work presents an interpretable machine learning (ML) system that uses composition- and physics-based descriptors to predict the bulk moduli of high-entropy alloys (HEAs) by permitting precise and computationally efficient bulk modulus prediction, as well as physically significant insights into descriptor–property connections.
Abstract
This work presents an interpretable machine learning (ML) system that uses composition- and physics-based descriptors to predict the bulk moduli of high-entropy alloys (HEAs). Extra Trees, Random Forest, Gradient Boosting, AdaBoost, and LightGBM are five ensemble ML algorithms that were systematically shaped and refined by hyperparameter fine-tuning. With a test R2 of about 0.852 and an RMSE and MAE of about 5.49 GPa and 1.5 GPa, respectively, Extra Tree outperformed the other optimized models, indicating good generalization capacity for untested HEA compositions. The computational efficiency results showed that LightGBM had the fastest prediction speed (~4.24 ms), whereas Extra Trees had the shortest training time (~17.3 s). The majority of the optimized models had statistically equal prediction performance (p > 0.05), according to statistical validation using paired t-test analysis, even though residual error distributions for the Extra Tree model established consistent and unbiased predictions. To enhance the interpretability of the model, SHAP-based explainable analysis was performed, which included SHAP importance, dependence, and waterfall plots. The SHAP results revealed that the primary determinants impacting bulk modulus behavior in HEAs were Zr content, mean electronegativity, Al content, bond strength, and melting-temperature-related parameters. The proposed framework enables the rapid identification and design of next-generation HEAs by permitting precise and computationally efficient bulk modulus prediction, as well as physically significant insights into descriptor–property connections.
Refractory high‐entropy alloys (RHEAs) offer mechanical performance at extreme temperatures, but their optimization is hindered by the vast compositional space and high cost of first‐principles calculations. This study establishes an interpretable machine‐learning benchmark for predicting nine elastic and mechanical properties of RHEAs. An EMTO‐CPA dataset comprising 2487 alloys is used to evaluate eight regression models: Gaussian process regression (GPR), support vector regression (SVR), shallow and deep neural networks, LightGBM, XGBoost, CatBoost, and Histogram‐based Gradient Boosting. Seven composition‐derived descriptors representing atomic size, electronic structure, and thermodynamic characteristics are employed. Hyperparameters are optimized using randomized search with five‐fold cross‐validation, while robustness is assessed through ten train–test evaluations using random seeds 42–51. Model performance is reported as mean ± standard deviation of MAE and R
2
. The results demonstrate target‐dependent performance: GPR performs particularly well for sws, SVR achieves favorable performance for several properties, and deep neural networks provide strong predictions for selected mechanical targets. SHAP analysis of optimized SVR models identifies valence electron concentration, average atomic radius, atomic size mismatch, and melting temperature as important contributors, whereas mixing entropy and mixing enthalpy generally show weaker contributions. Overall, the benchmark provides a reproducible framework for RHEA prediction and supports screening and alloy design.
Amer Almahmoud, A. Obeidat· Advanced Theory and Simulati...· 0 citations
High-entropy alloys (HEAs) exhibit exceptional stability in extreme environments, yet their expansive design space presents a “curse of dimensionality” for traditional discovery methods. While machine learning (ML) offers a data-driven paradigm for material screening, the scarcity of experimental data often results in overfitting and limited physical interpretability. To address these challenges, this study proposes a hybrid physics-informed machine learning (Hybrid PIML) framework for accelerated hardness prediction. By integrating classical solid solution strengthening theory with a residual learning artificial neural network (ANN), the model explicitly embeds the physical coupling of shear modulus and lattice distortion (G·δr2/3) as prior knowledge. This approach ensures predictions adhere to metallurgical principles while significantly outperforming benchmark algorithms, achieving a coefficient of determination (R2) of 0.976 and reducing the root mean square error (RMSE) by approximately 43%. SHapley Additive exPlanations (SHAP) analysis confirms that physics-enhanced features dominate the decision-making process, validating the model’s internalization of strengthening mechanisms. Furthermore, the research elucidates a phase-dependent non-linear correlation between hardness and yield strength, correcting the failure of the classical Tabor formula in work-hardening face-centered cubic (FCC) alloys. Finally, a high-throughput virtual screening funnel based on this framework successfully identified optimized non-equiatomic candidates within the refractory Co-Cr-Ti-Mo-W system. This work establishes a precise, physically consistent pathway for inverse material design under data-constrained conditions.
Ao-Yuan Gao, Yi-Xiang Yan, Qiu-Ling Tao et al.· Journal of Materials Informa...· 0 citations
The Northeast Materials Database is leveraged to develop machine learning models that predict magnetic materials with targeted Curie temperatures from composition-derived descriptors rooted in molecular-level elemental properties, supplemented by a small set of coarse crystal-system and structure-family indicators.
F. Uçar, Nida Katı· Scientific Reports· 0 citations
High-entropy alloys (HEAs) provide a broad compositional space for developing structural materials with balanced phase stability and mechanical performance. However, reliable mechanical-property prediction remains challenging because alloy chemistry, phase constitution, processing state, and model uncertainty are strongly coupled.
We developed a phase-stability-guided explainable machine learning framework using a curated database comprising 541 phase-labelled HEAs, 263 hardness-labelled records, and 214 yield-strength-labelled records. Hierarchical physical and processing descriptors were combined with leakage-controlled phase-probability features generated through composition-level nested cross-fitting. XGBoost and CatBoost models were used for property prediction, bootstrap ensembles for uncertainty quantification, and SHAP and accumulated local effects for model interpretation. A total of 300,000 virtual candidates were screened, followed by CALPHAD-assisted assessment and experimental validation of three representative alloys.
The phase classifier achieved an overall accuracy of 0.923. The phase-guided property models achieved R² values of 0.881 for hardness and 0.856 for yield strength. The experimentally measured dominant phases, hardness values, and compressive yield strengths of three representative candidates were generally consistent with the model predictions, with moderate experimental deviations.
The proposed framework integrates physical descriptors, probabilistic phase information, model explainability, and uncertainty-aware screening, providing an interpretable and experimentally supported strategy for prioritizing promising HEA compositions before broader experimental optimization.
An explainable machine-learning framework was developed for dielectric constant prediction using 52,168 crystalline materials extracted from the Joint Automated Repository for Various Integrated Simulations (JARVIS-DFT) database, demonstrating the complementary roles of electronic structure and elemental chemistry.
D. Pundhir, Ashok Kumar· Applied Physics A· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.