Jul 2026· Journal of Chemical Information and Modeling· Vol 66, pp. 7996-8007· 1 citation· 34 references
MedicineComputer Science
TL;DR
This work compares 26 feature importance pipelines spanning data-driven, model-based, and formula-based analyses on a metal-support interaction data set anchored by an explicit SISSO equation, and examines whether the same qualitative behavior recurs in high-entropy-alloy and halide perovskite data sets.
Abstract
While feature importance analysis in chemical and materials machine learning can often be sensitive to both the predictive model and the attribution rule, the robustness of these rankings is rarely quantified before they are used to gain physical insights. Here, we compare 26 feature importance pipelines spanning data-driven, model-based, and formula-based analyses on a metal-support interaction data set anchored by an explicit SISSO equation, and we examine whether the same qualitative behavior recurs in high-entropy-alloy and halide perovskite data sets. Across the three benchmarks, we observe high intrafamily agreement but substantial interfamily variance. While a small subset of features remains stable across multiple families, several midranked features are highly family dependent, with their apparent importance shifting according to the underlying modeling assumptions. To ensure robust interpretability, we recommend that feature importance be reported by method family or correlation-based clusters, supplemented by resampling intervals.
Predicting thermal stability during handling and storage is essential for the design of safe and reliable energetic materials. However, experimental measurements vary significantly across laboratories due to differences in protocols and analysis methods, making it difficult to train reliable predictive models. We addre...
M. Davis, R. Ullberg, J. Schroeder et al.· 0 citations
Feature Selection (FS) is an essential data preprocessing technique aimed at identifying the most relevant features for predictive models. However, traditional FS approaches often struggle to capture complex interactions in real-world datasets, limiting their ability to fully support high-performing Machine Learning (M...
Covalent organic frameworks (COFs) are highly ordered, porous organic materials whose reticular construction from tailored nodes and linkers enables atomic-level control over structure and function. The design space of COFs is vast with virtually unlimited combinations of nodes, linkers, and functional groups. Interpre...
Alathea E. Davies, O. Adesina, Isabella M. Valdez et al.· Journal of Chemical Theory a...· 1 citation
The Northeast Materials Database is leveraged to develop machine learning models that predict magnetic materials with targeted Curie temperatures from composition-derived descriptors rooted in molecular-level elemental properties, supplemented by a small set of coarse crystal-system and structure-family indicators.
F. Uçar, Nida Katı· Scientific Reports· 0 citations
Machine learning is increasingly used to learn structure property relationships from spectroscopic and diffraction data, yet its adoption in materials discovery is often limited by poor interpretability of model predictions. Although attention mechanisms are frequently treated as inherently explainable, unregularized a...
Aditya Raghavan, Utkarsh Pratiush, Dalton A. Pearl et al.· 0 citations
Feature importance is widely used in machine learning to assess the contribution of individual predictors to model performance and support the interpretation of model behaviour. However, it remains unclear whether features identified as highly important are also those to which a model is most sensitive when their value...
Simran Sharma, Maheshkumar Mulani· International Journal For Mu...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.