Skip to content

XRF-to-Optical Field-of-View Localization with Vision Language Models

Aug 2026 · 0 citations · 39 references
Computer Science

TL;DR

This paper evaluates training-free vision language model (VLM) localization on two datasets representing same-section high-correspondence and adjacent-section low-correspondence imaging and tests unconstrained and metadata-constrained search and VLMs with geometric controls, classical template matching, and two alternative training-free approaches.

Abstract

Registering images acquired with different microscopy modalities is essential for relating complementary measurements of the same specimen. In correlative X-ray fluorescence (XRF) and optical microscopy, the XRF map often covers only a small region of an optical image acquired from the same or an adjacent tissue section. Field-of-view (FOV) localization is necessary but can be difficult when appearance and structure differ across modalities. Here we evaluate training-free vision language model (VLM) localization on two datasets representing same-section high-correspondence and adjacent-section low-correspondence imaging. We test unconstrained and metadata-constrained search and compare VLMs with geometric controls, classical template matching, and two alternative training-free approaches (DINOv2 and multiGradICON). Direct VLM prompting produced content-dependent spatial signals but was not reliable alone. Classical matching was most accurate when cross-modal structure was preserved but failed in the low-correspondence collection. A proposal-and-verify workflow used repeated VLM predictions as candidates and image-based similarity to select the final location. This workflow recovered useful localization in the low-correspondence regime.

View source

Similar papers

Review Aug 2026

Acquisition Geometry-Assisted Whole-Group Localization of X-ray Fluorescence Maps in Optical Microscopy Images

X-ray fluorescence (XRF) microscopy maps elemental distributions, while optical microscopy can provide complementary morphological context. Localizing XRF fields of view (FOVs) in optical images is difficult because the two modalities differ in contrast mechanism and resolution. Most current workflows place each XRF tile independently, even when acquisition metadata already record the tiles'relative scan positions. This study formalizes XRF tile-group localization, in which one optical-frame placement is estimated for the whole group, constrained by acquisition geometry and quantified using group intersection-over-union (GroupIoU). In a controlled case study, independent localization failed with GroupIoU 0.000, whereas group localization achieved 0.931. Replacing the normalized cross-correlation (NCC) metric with mutual information (MI) gave nearly identical results, showing that the outcome is not specific to one local similarity metric. In another multiscale case study, using a coarse XRF survey scan to connect the fine-scale tile group to the optical image increased mean GroupIoU from 0.694 to 0.856. These case studies support using acquisition geometry as an explicit constraint when localizing related XRF tiles.

Xiangyu Yin, T. Paunesku, Letonia Copeland-Hardin et al. · 0 citations
Aug 2026

Multiple Modalities Image Matching With Large-Scale Pre-Training.

Image matching, which aims to identify corresponding pixel locations between images, is crucial in a wide range of scientific disciplines, aiding in image registration, fusion, and analysis. However, when dealing with images captured under different imaging modalities that result in significant appearance changes, the performance of learning-based image matching algorithms often deteriorates due to the scarcity of annotated cross-modal training data. This limitation hinders applications in various fields that rely on multiple image modalities to obtain complementary information. To address this challenge, we propose a large-scale pre-training framework that utilizes synthetic cross-modal training signals, incorporating diverse data from various sources, to teach models to recognize and match fundamental structures across images. This capability is transferable to real-world, unseen cross-modality image matching tasks. Our key finding is that the matching model trained with our framework generalizes effectively across more than eight unseen cross-modality registration tasks using the same set of network weights, substantially outperforming existing generalizable methods and achieving competitive or superior performance compared to specialized models on several tasks.

Xingyi He He, Hao Yu, Sida Peng et al. · 0 citations
Conference Jul 2026

Structured CT Imaging Artefact Assessment using Vision-Language Models

X-ray computed tomography (CT) is one of today’s most critical imaging modalities, with a wide range of applications spanning from medical diagnosis to industrial inspection. CT images can be severely affected by physically induced degradations such as low-dose noise and beam hardening, which compromise image quality and diagnostic accuracy. Existing AI-based approaches largely treat this problem as a pure classification task, and a system that explains the physical mechanisms of artefacts and provides actionable recommendations to the user has not been systematically addressed. In this study, we propose a vision-language model (VLM) based pipeline that detects CT artefacts, explains their physical mechanisms in natural language, and generates structured, actionable recommendations. LLaVA-1.5-7B and Qwen2-VL-7B models were fine-tuned using QLoRA on the 2DeteCT dataset; following fine-tuning, LLaVA-1.5-7B achieved 99.8% accuracy while Qwen2-VL-7B reached 86.9%. The results demonstrate the effectiveness of domain adaptation for structured artefact assessment.

Reyhan Hosavci, Raziye Kübra Kumrular · 0 citations
Open access Aug 2026

A Correlative Microscopy Dataset for Multimodal Data Fusion and Image Matching in Materials Science

Gaining understanding of process-structure-property relationships in materials at a mechanistic level relies on correlative microscopy workflows. These workflows, in turn, fundamentally depend on image matching, i.e., a computer vision task with the objective of finding point correspondences in pairs of images. Matching models are difficult to evaluate quantitatively in the materials field due to a shortage of representative benchmark datasets. Nonetheless, prior research indicates that traditional rule-based image matching techniques such as the surface-invariant feature transform (SIFT) currently fall short on such matching tasks. We present a dataset for cross-modal image matching and data fusion in the materials microscopy domain, which we coin AmalgaMatch , to support model benchmarking and fine-tuning efforts. All images are micrographs captured using the most widely applied imaging techniques in materials science including light-optical, scanning electron, and transmission electron microscopy, as well as electron backscatter diffraction (EBSD). Therein, various detectors and imaging modes are employed to capture micrographs of diverse materials. While the majority of images are raw images, some underwent typical processing routes using digital image correlation or EBSD indexing. Common regions in image pairs are populated with hand-annotated keypoint correspondences. While mutual information is limited in cross-modal, multi-scale image pairs, we relied on characteristic defects such as dislocations, grain boundaries, triple junctions, inclusions, pores or topographic features for annotation. Furthermore, the dataset is divided into groups, defined by distinct registration use cases, and further into subsets, defined by the imaged material. The dataset covers many typical use cases for image matching in materials science, including slip partitioning, dislocation characterization, and surface fractography. In total, it comprises 6 groups and 19 subsets with 35 scenes and 187 annotated image pairs to support autonomous multimodal materials data fusion. For each image, we provide structured metadata to facilitate training of hybrid matching models which process textual alongside image-based inputs to improve the matching quality and robustness. A formal ontological model for correlative microscopy and image matching processes is proposed to express image contents, relationships, and transformations through knowledge graphs and to enable aligning with FAIR data principles.

Ali Riza Durmaz, James D. Lamb, Kamilla Zaripova et al. · 1 citation
Review Open access Jul 2026

Review of image segmentation techniques for biomedical micro-CT: from laboratory absorption imaging to synchrotron phase-contrast

X-ray micro-computed tomography (μCT) is widely used in biomedical research for non-destructive, high-resolution imaging. Synchrotron Radiation Phase-Contrast μCT (SR-PCI-μCT) further enhances image quality with higher Signal to Noise Ratio (SNR), improved contrast, and faster acquisition. However, segmenting SR-PCI-μCT images remains challenging due to their heterogeneous image property and limited training data. Additionally, the increasing throughput of synchrotron facilities demands robust, efficient segmentation methods. This paper reviews a list of recent segmentation approaches in biomedical μCT. Traditional methods remain simple and effective for simple, high-contrast structures but require extensive tuning and generalize poorly to complex, low-contrast soft tissues. Data-driven models provide higher accuracy and robustness yet rely heavily on large expert-annotated datasets, limiting reproducibility and cross-dataset adaptability. Recent advance on vision transformers have shifted the paradigm from task-specified to more domain-specified segmentation, though these techniques are still evolving and require adaptation for SR-PCI-μCT studies. This survey provides the first bi-modality review covering both laboratory and synchrotron biomedical μCT segmentation. It consolidates recent segmentation methodologies, identifies major trends in deep-learning techniques, and highlights current limitations across SR-PCI-μCT. Additionally, it outlines open challenges to guide future research and practical advancements in biomedical μCT segmentation.

Hao Song, Ning Zhu · 0 citations
Preprint Aug 2026

Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging

Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tasks like glioma grading. While visual-language models (VLMs) show promise in cross-modal understanding, existing methods focus mainly on 2D image modeling, neglecting direct perception of 3D volumetric space. Although 3D VLMs have been proposed for report generation and feature alignment in 3D CT imaging, mpMRI applications demand collaborative inference across multiple imaging modalities-a requirement unmet by current solutions. To address this, we introduce Mr3D-VL, a dedicated visual-language foundation model for multi-parametric 3D MRI. With 4 billion parameters, it employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding for dual modality-spatial integration. Its cross-modal projection layer uses a multi-resolution feature implantation strategy to enhance feature perception across resolutions. Experimental results show significant improvements over existing 4B/7B/30B domain-specific and general-purpose models in text generation tasks, achieving a BERTScore of 0.856 for report generation, with question-answering accuracy at 0.713 and multiple-choice accuracy at 0.912.

Zhi Qiao, Xintong Wu, Yichu He et al. · 0 citations

Related blog posts

Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.