Skip to content
Review Open access

Multimodal Machine Learning for Data-Driven Materials Selection: A Review of Foundations, Advances, and Future Directions

2026 · IEEE Access · Vol 14, pp. 126239-126266 · 0 citations · 95 references

TL;DR

This review aims to offer a conceptual framework and a solid reference for building intelligent multimodal material selection systems and presents a forward-looking research agenda covering self-supervised learning, knowledge-enhanced models, and interpretable human–AI collaboration.

Abstract

Data-driven material selection is progressively changing how materials are evaluated in engineering, manufacturing, and product design. With the growing diversity of heterogeneous data sources—ranging from microscopic images and physicochemical properties to simulation outputs and textual data—multimodal machine learning (MML) has become a key technology for fusing different types of information and supporting reliable multi-objective decision-making. However, despite increasing research interest, there is still no unified and conceptually structured review on the application of MML methods in data-driven material selection. This paper provides a thorough and structured overview of the field, organized using a novel multidimensional classification framework based on data types, integration levels, learning paradigms, and decision-making tasks. Guided by this framework, we critically evaluate the strengths, limitations, and general applicability of existing methods, and identify current trends and research gaps. Beyond qualitative synthesis, we quantify the cited corpus (N = 94) by publication year, method family, and application area, and collate an empirical fusion-evidence table that reports each study’s quantified gain together with its boundary conditions. We further analyze the main challenges and present a forward-looking research agenda covering self-supervised learning, knowledge-enhanced models, and interpretable human–AI collaboration. This review aims to offer a conceptual framework and a solid reference for building intelligent multimodal material selection systems.

Read PDF

Similar papers

Open access Jul 2026

Generative and multimodal AI for materials prediction and design: progress, challenges, and perspectives

A materials property hierarchy is introduced, from intrinsic, composition-determined properties to extrinsic, processing-dependent performance, to clarify deployment constraints and distinguish structural, physical and deployment novelty.

Xianyuan Liu, Charles Anjah, Benjamin E. Jolly et al. · 0 citations
Review Open access Jul 2026

Bridging modalities: a Survey and Taxonomy of Automated Multimodal Machine Learning (MMAutoML)

Automated machine learning (AutoML) reduces the cost of developing high-performing models by automating pipeline design, model selection, and hyperparameter optimization. As data sets increasingly combine heterogeneous sources (e.g., text, images, audio, and structured/tabular records), AutoML is extending to multimodal settings, where performance depends on representation learning and cross-modal fusion as well as model search. This survey reviews automated multimodal machine learning (MMAutoML), distinguishing core systems that automate both representation and fusion decisions from near-core systems, partial multimodal AutoML tools, and adjacent multimodal ML infrastructure. We organize multimodal learning around early, late, and hybrid fusion paradigms and discuss trade-offs in robustness, interpretability, and deployment. We then curate and compare representative open-source and commercial MMAutoML frameworks, summarize supported modalities, and characterize the ecosystem via time-evolution and similarity-based groupings. Finally, we overview applications in healthcare, autonomous systems, finance, and e-commerce, and highlight open challenges in modality alignment, missing or degraded inputs, efficiency and scalability, reproducibility, and governance (privacy, bias, and monitoring). We conclude with practical guidance and research directions toward reliable, end-to-end MMAutoML.

Blaž Škrlj, Aljaž Osojnik · 0 citations
Review Open access Aug 2026

Data-Driven Materials Science for Energy-Sustainable Applications.

Materials science is underpinned by structure-property relationships that govern the function of a material. These relationships can be encoded into algorithms and integrated into machine-learning models that enable the prediction of materials and their cognate properties. However, machine-learning models are largely being trained on computed data owing to a worldwide shortage of real-world (experimental) datasets. This review describes how to capture and collate experimental data from scientific literature using artificial-intelligence (AI) methods to produce materials-domain-specific datasets or language models. Their application in AI-driven enquiries that facilitate progress in energy-sustainable materials science is then illustrated via six case studies that cover: training machine-learning models, data-driven materials discovery, optimizing manufacturing processes, mapping phases of materials, forecasting materials-centric research trends, and classifying types of materials using automated prompt engineering. The future of materials-domain-specific datasets, language models, and decision-making workflows using AI agents is then envisioned for the energy sector. The intrinsic challenges of accessing historical dark data in materials science are then described and contrasted with timely opportunities for leveraging massive amounts of experimental data from laboratories in going forwards; by exploiting electronic-lab notebooks, high-throughput experiments, and digital-twin technologies. These opportunities are illustrated for energy-sustainable materials science, especially the photovoltaic and battery industries.

Jacqueline M. Cole · 0 citations
Jul 2026

Data-Driven Material Design: Harnessing High-Throughput Simulations and AI

This work will present the current work on inverse material design, where AI methods—particularly generative pretrained transformers—are used to predict new material candidates based on desired properties, pushing the boundaries of materials innovation.

I. Gonzales, R. Ullberg, Andrew H Salij et al. · 0 citations
Open access Aug 2026

Improving Composite Materials with Machine Learning: A Predictive Approach

Machine learning (ML) has become an important technology in the field of composite materials, offering efficientmethodsformaterialcharacterization, damageassessment, propertyprediction, anddesignoptimization. Conventional experimental techniques and physics-based simulations for predicting the complex behavior of composites can require considerable time, cost, and computational resources. In contrast, ML provides a data-driven approach capable of identifying hidden patterns, establishing relationships between variables, and generating accurate predictions. Algorithms such as Support Vector Regression (SVR), Random Forest Regression, and Linear Regression can process information related to material composition, filler characteristics, curing conditions, and environmental factors. This study examines the application of ML techniques to composite materials, particularly for predicting fracture toughness, characterizing damage, and optimizing mechanical properties. Correlation analysis reveals significant relationships between fracture toughness and important input parameters, demonstrating the usefulness of data-driven approaches in material development. Among the investigated techniques, Random Forest Regression demonstrates strong predictive capability and effectively represents complex material behavior. However, challenges including limited data availability, model interpretability, and generalization remain. Reliable ML performance depends on high-quality datasets, suitable feature selection, and effective model tuning. Overfitting can also affect prediction accuracy, making cross-validation and hyperparameter optimization essential. Integrating ML into composite material research can reduce experimental requirements, lower development costs, improve manufacturing efficiency, and accelerate the development of high-performance materials. Future research should emphasize larger datasets, explainable ML, and hybrid approaches combining ML with physics-based simulations.

Periyasamy Chitra · 0 citations
Open access Aug 2026

Mixture-of-Experts Learning for Mixture-Response Interpretation and Screening of PE-ECC

Highlights A PE-ECC database comprises 383 material-level records from 90 literature sources. The MoE model predicts four PE-ECC properties with R2 values of 0.950–0.971. SHAP, ALE and response maps distinguish strength trends from tensile deformation. Binder, W/B, S/B and fiber variables show property-specific relationships. Support-filtered screening yields database-supported candidates for laboratory validation. Abstract Featuring considerable tensile ductility and multiple cracking behavior, polyethylene fiber-reinforced engineered cementitious composites (PE-ECCs) are promising cement-based materials for engineering construction. However, establishing accurate design models for evaluating the mechanical properties of PE-ECC is a challenging task owing to the complex material components. This study presents an interpretable data-driven framework for predicting the mechanical properties of PE-ECC using mixture-of-experts (MoE) learning. A database comprising 383 deduplicated material-level records from 90 verified literature sources was compiled for modeling the compressive strength, ultimate tensile strain, ultimate tensile strength and first-cracking tensile strength of PE-ECC. An MoE prediction model was developed by integrating XGBoost, LightGBM, CatBoost, WDBPANN and TabPFN through out-of-fold stacking and learned gating. The model achieved coefficient of determination (R2) values of 0.971, 0.950, 0.970 and 0.954 for the four mechanical properties, respectively. Shapley additive explanations (SHAP), accumulated local effects (ALE) and response maps were used to examine the fitted nonlinear associations between the reported mixture variables and each target property. Based on these relationships, support-filtered virtual screening was conducted within the database-supported design space to identify candidate mixtures for subsequent experimental verification. The framework links target-specific prediction with mixture-response interpretation and confines screening to regions supported by reported PE-ECC mixtures.

Yujie Wang, Lingzhi Li · 0 citations