A Correlative Microscopy Dataset for Multimodal Data Fusion and Image Matching in Materials Science – Scientific Data
Materials scientists increasingly rely on correlative microscopy to connect how a material is made, what its internal structure looks like, and how it ultimately performs. That workflow often requires researchers to compare and align images taken from very different instruments. In practice, this means solving an image-matching problem: finding the same physical points or features across two separate micrographs.
That sounds straightforward, but in materials science it is anything but. Images may come from different microscopes, different detectors, different scales, and even different physical contrast mechanisms. A crack, grain boundary, pore, or dislocation may appear dramatically different depending on the imaging method. This makes accurate image registration one of the major barriers to automated multimodal data fusion.
The newly introduced dataset, called AmalgaMatch, is designed to tackle exactly that problem. It provides a benchmark for cross-modal image matching in materials microscopy and aims to support both model evaluation and model fine-tuning. In a field where representative public datasets have been scarce, that is a notable step forward.
Traditional image-matching approaches, especially rule-based methods such as SIFT (Scale-Invariant Feature Transform), have long been used in computer vision. But prior work suggests that these methods often struggle with materials microscopy data, particularly when matching images across different imaging modalities. AmalgaMatch addresses this gap by offering a curated set of annotated micrographs tailored to real materials-science use cases.
The dataset draws from some of the most widely used imaging techniques in the field. These include light optical microscopy, scanning electron microscopy, transmission electron microscopy, and electron backscatter diffraction (EBSD). It also spans multiple detectors and imaging modes, covering a diverse range of materials and experimental conditions.
Most of the images are raw micrographs, preserving the complexity researchers encounter in practice. Some, however, include standard processing outputs such as digital image correlation results or EBSD indexing. That mix is important because real-world correlative workflows rarely involve perfectly uniform data. Instead, they combine images and derived maps produced at different stages of analysis.
One of the dataset’s most valuable features is its use of hand-annotated keypoint correspondences. For each image pair, common regions were manually marked to identify matching points. Because cross-modal and multi-scale images often share limited direct visual similarity, the annotations focus on distinctive material features such as dislocations, grain boundaries, triple junctions, inclusions, pores, and topographic structures. These are the kinds of landmarks that human experts actually use when aligning microscopy data.
AmalgaMatch is organized in a way that reflects practical scientific workflows. It is divided into six groups, each representing a distinct registration use case, and then further split into 19 subsets based on the material being studied. Altogether, the dataset contains 35 scenes and 187 annotated image pairs.
That scale is significant not just because of the number of images, but because of the diversity of matching scenarios it covers. The benchmark includes applications such as slip partitioning, dislocation characterization, and surface fractography—all areas where multimodal correlation can reveal important process-structure-property relationships.
Another forward-looking aspect of the release is its inclusion of structured metadata for every image. This opens the door to hybrid AI systems that do not rely on pixels alone, but also use text-based context to improve image matching. For example, a model could potentially incorporate information about microscope type, detector mode, material class, or sample preparation when deciding how two images should align. That could make future matching systems more robust, especially in difficult low-overlap or low-similarity scenarios.
The authors also propose a formal ontological model for correlative microscopy and image-matching workflows. In simpler terms, this is a framework for describing image content, transformations, and relationships in a structured, machine-readable way. Such a model could help build knowledge graphs around microscopy datasets and improve interoperability across labs, software tools, and repositories.
Just as importantly, this semantic layer supports alignment with FAIR data principles—making data more findable, accessible, interoperable, and reusable. For a field that increasingly depends on machine learning and large-scale data integration, that is more than a documentation improvement; it is infrastructure.
AmalgaMatch arrives at a time when materials informatics is pushing further into multimodal AI. If machine learning systems are to understand complex materials behavior, they must be able to connect information from multiple imaging sources reliably. That starts with better benchmarks.
By providing annotated cross-modal microscopy pairs, realistic use cases, and metadata-rich structure, AmalgaMatch gives researchers a much-needed platform for testing and improving matching algorithms. For the materials science community, it may become an important building block in the larger effort to automate correlative microscopy and unlock deeper insights into how materials work.