Skip to content

Cross-modal learning for SAR target recognition using optical vision foundation models

Sep 2026 · Artificial Intelligence for Security and Defence Applications IV · pp. 66 · 0 citations · 22 references
Computer Science Engineering

TL;DR

This work proposes a cross-modal EO to SAR prototype alignment framework in which a frozen EO encoder, based on a DINOv3 vision foundation model, is used to construct class level optical prototypes without requiring strict EO/SAR pairs.

Abstract

Synthetic Aperture Radar (SAR) is an important modality in a wide range of imaging applications due to its versatile, long range and near all weather operating capabilities. However, Automatic Target Recognition (ATR) remains a challenging problem due to limited labelled data, the strong speckle in SAR images and the significant domain gap between SAR and more abundant optical imagery. In contrast, electro-optical (EO) imagery benefits from massive datasets, clearer visual structure and powerful foundation models. In this work, we investigate how vision foundation models trained on optical data can provide class level supervision for SAR classification. We propose a cross-modal EO to SAR prototype alignment framework in which a frozen EO encoder, based on a DINOv3 vision foundation model, is used to construct class level optical prototypes without requiring strict EO/SAR pairs. A SAR model is then trained to classify SAR images while aligning its embeddings to the corresponding EO class prototype. At inference time, the SAR model operates independently, without access to optical imagery. We evaluate our approach on the UNICORNv2 dataset, an EO and SAR dataset of civilian vehicles with heavily speckled images and severe class imbalance. EO prototype alignment improves SAR classification accuracy over frozen DINOv3, SAR only finetuning and unpaired distribution alignment baselines, and t-SNE visualizations provide qualitative evidence of clearer separation among classes in the trained SAR embedding space. These results suggest that optical vision foundation models, despite being trained on visible spectrum imagery, provide transferable information for SAR image classification, offering a practical method for using large scale pretrained vision foundation models across challenging sensing modalities.

Read PDF

Similar papers

2026

Bridging Optical Appearance and SAR Scattering With Vision-Language Prototypes for Zero-Shot Target Recognition

Supervised synthetic aperture radar (SAR) automatic target recognition (ATR) methods rely on a closed-set assumption and struggle to recognize unseen target categories without labeled SAR samples. Zero-shot learning (ZSL) offers a promising solution by transferring knowledge from seen classes and external semantic prio...

Rui Zhu, Tian-Wen Zhang, Xiao-Ling Zhang et al. · 1 citation
Open access Oct 2026

A Three-Branch SAR Target-Recognition Network via Multi-Physical-Dimensional Collaboration

Conventional deep-learning methods for synthetic aperture radar (SAR) target recognition predominantly rely on image-domain visual features, with insufficient integration of the physical principles governing SAR imaging. This limitation leads to unsatisfactory recognition accuracy and poor generalization performance un...

Chen Dong, Shuang-Xi Zhang, Hong-Tao Zhan et al. · 0 citations
Open access Aug 2026

Context-Guided Discrimination Feature Learning in Color Space for Aircraft Detection in SAR Images

Aircraft often manifest as highly aspect/pose-sensitive, disjointed blobs made of pixels with fluctuating levels of brightness in classic grayscale SAR images, which makes the aircraft annotation task challenging even for human experts. As more and more high-resolution colorized SAR images acquired by the latest commer...

Yu Zhang, Zhe Geng, Lu-Jian Yao et al. · 0 citations
Preprint Sep 2026

Improved Automatic Target Recognition in Synthetic Aperture Sonar Imagery Using Large Deep Neural Networks

Automatic Target Recognition (ATR) in Synthetic Aperture Sonar (SAS) is a task largely dominated by deep neural networks (DNNs). Most SAS-ATR models use convolutional neural network (CNN) architectures whereas transformer-based architectures have had much less representation in the literature despite being state of the...

C. Moore, A. Hurt, Jordan M. Malof · 0 citations
Preprint Sep 2026

GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation

Paired synthetic aperture radar (SAR) and electro-optical (EO) imagery is increasingly available across sensors, resolutions, and geographic regions. Yet existing SAR-to-EO image translation (SET) methods are typically trained on a single, limited-scale dataset, producing models specialized to particular sensing condit...

Jeonghyeok Do, Munchurl Kim · 0 citations
Sep 2026

Deep Learning-Based Classification of Military Aircraft Types from Satellite Imagery

Automatic classification of military aircraft in satellite imagery is a challenging problem with a high risk of error due to the limited pixel area of targets, variations in image resolution and illumination conditions, background complexity, and strong visual similarity among classes. In this study, a deep learning ap...

Fahrettin Varlik, Cüneyt Özdemir · 0 citations

Related blog posts

Microsoft Research Blog Aug 11, 2026

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.