Skip to content
Open access

A Correlative Microscopy Dataset for Multimodal Data Fusion and Image Matching in Materials Science

Aug 2026 · Scientific Data · 1 citation

Abstract

Gaining understanding of process-structure-property relationships in materials at a mechanistic level relies on correlative microscopy workflows. These workflows, in turn, fundamentally depend on image matching, i.e., a computer vision task with the objective of finding point correspondences in pairs of images. Matching models are difficult to evaluate quantitatively in the materials field due to a shortage of representative benchmark datasets. Nonetheless, prior research indicates that traditional rule-based image matching techniques such as the surface-invariant feature transform (SIFT) currently fall short on such matching tasks. We present a dataset for cross-modal image matching and data fusion in the materials microscopy domain, which we coin AmalgaMatch , to support model benchmarking and fine-tuning efforts. All images are micrographs captured using the most widely applied imaging techniques in materials science including light-optical, scanning electron, and transmission electron microscopy, as well as electron backscatter diffraction (EBSD). Therein, various detectors and imaging modes are employed to capture micrographs of diverse materials. While the majority of images are raw images, some underwent typical processing routes using digital image correlation or EBSD indexing. Common regions in image pairs are populated with hand-annotated keypoint correspondences. While mutual information is limited in cross-modal, multi-scale image pairs, we relied on characteristic defects such as dislocations, grain boundaries, triple junctions, inclusions, pores or topographic features for annotation. Furthermore, the dataset is divided into groups, defined by distinct registration use cases, and further into subsets, defined by the imaged material. The dataset covers many typical use cases for image matching in materials science, including slip partitioning, dislocation characterization, and surface fractography. In total, it comprises 6 groups and 19 subsets with 35 scenes and 187 annotated image pairs to support autonomous multimodal materials data fusion. For each image, we provide structured metadata to facilitate training of hybrid matching models which process textual alongside image-based inputs to improve the matching quality and robustness. A formal ontological model for correlative microscopy and image matching processes is proposed to express image contents, relationships, and transformations through knowledge graphs and to enable aligning with FAIR data principles.

Read PDF