Skip to content
Preprint

CrossFeat: Bridging Imaging Modalities in Feature Descriptor Space

Aug 2026 · 0 citations · 55 references
Computer Science

TL;DR

This work proposes CrossFeat, a framework that enables an existing monomodal descriptor to operate across modalities, and introduces a geometry-appearance disentanglement such that only appearance is altered while the geometric properties are preserved.

Abstract

Most advances in keypoint descriptions address monomodal settings, where image variations arise from viewpoint, illumination, or contrast changes. Multimodal scenarios involve images produced by fundamentally different sensing processes, such as multispectral imaging, RGB-depth, satellite imagery, or medical imaging, causing the same structures to appear differently. A common solution to cross-modal description is to train descriptors for each modality pair, which requires retraining whenever the modalities change, or to train large models, which incur a significant increase in runtime. Instead, we propose CrossFeat, a framework that enables an existing monomodal descriptor to operate across modalities. Our method learns a crossing function in descriptor space that maps features from one modality to a representation compatible with another. To preserve the structural information captured by the original descriptor, CrossFeat introduces a geometry-appearance disentanglement such that only appearance is altered while the geometric properties are preserved. Experiments across multiple domains and datasets demonstrate improved performance in multimodal matching.

View source

Similar papers

Aug 2026

Multiple Modalities Image Matching With Large-Scale Pre-Training.

A large-scale pre-training framework that utilizes synthetic cross-modal training signals, incorporating diverse data from various sources, to teach models to recognize and match fundamental structures across images, which generalizes effectively across more than eight unseen cross-modality registration tasks.

Xingyi He He, Hao Yu, Si-Da Peng et al. · 0 citations
Preprint Sep 2026

Radiation, Rotation and Scale Invariant Feature Descriptor for Multimodal Image Matching

A radiation, rotation, and scale invariant (RRSI) feature descriptor that enables feature encoding, interaction, and fusion across intra-modal, dual-head sampled, and inter-modal regions, and introduces a bidirectional cross-modal generative reconstruction constraint during training.

Yuan-Xin Ye, Teng-Feng Tang, Tao Peng et al. · 0 citations
Preprint Sep 2026

S2T-Unet: A Structure-to-Style Framework for Inter-Modality MRI Translation

Inter-modality MRI translation aims to synthesize missing MRI modalities from available acquisitions, reducing the need for additional scanning while preserving clinically relevant anatomical information. However, existing image translation methods often learn intensity mappings without explicitly separating modality-i...

Yi-Chao Liu · 0 citations
#machine learning Preprint Sep 2026

Binding Multiple Modalities via Multimodal Wasserstein Barycenter

Multimodal learning beyond two modalities commonly leverages a specific modality (e.g., text) to bind other modalities. However, how to establish a more balanced representation space that approximates shared semantics while respecting the holistic geometry of $n$-modal data remains challenging. In this work, we present...

Xiao-Le Tang, Jia-Yi Xu, Xiang Gu et al. · 0 citations
Conference Sep 2026

Understanding Domain-Shift Immunity in Deep Deformable Registration

It is shown that domain-shift immunity is an inherent, largely architecture-agnostic property of deep de-formable registration when trained with a robust pipeline and offered a principled explanation for the cross-domain generalizability of deep registration networks.

Mingzhen Shao, Sarang C. Joshi · 0 citations
Open access Sep 2026

MCF-Net: Multi-Modal Cross-Attention Fusion Network with Difference Convolution for Object Detection

Object detection based on visible-light images faces significant challenges under complex illumination and environmental conditions. Introducing infrared or depth images as complementary modalities can effectively enhance detection performance in such scenarios. However, most existing methods only support bi-modal conf...

Ming-Chun Li, Xin-Ran Wu, Rui Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.