Sep 2026· Artificial Intelligence for Security and Defence Applications IV· pp. 66· 0 citations· 22 references
Computer ScienceEngineering
TL;DR
This work proposes a cross-modal EO to SAR prototype alignment framework in which a frozen EO encoder, based on a DINOv3 vision foundation model, is used to construct class level optical prototypes without requiring strict EO/SAR pairs.
Abstract
Synthetic Aperture Radar (SAR) is an important modality in a wide range of imaging applications due to its versatile, long range and near all weather operating capabilities. However, Automatic Target Recognition (ATR) remains a challenging problem due to limited labelled data, the strong speckle in SAR images and the significant domain gap between SAR and more abundant optical imagery. In contrast, electro-optical (EO) imagery benefits from massive datasets, clearer visual structure and powerful foundation models. In this work, we investigate how vision foundation models trained on optical data can provide class level supervision for SAR classification. We propose a cross-modal EO to SAR prototype alignment framework in which a frozen EO encoder, based on a DINOv3 vision foundation model, is used to construct class level optical prototypes without requiring strict EO/SAR pairs. A SAR model is then trained to classify SAR images while aligning its embeddings to the corresponding EO class prototype. At inference time, the SAR model operates independently, without access to optical imagery. We evaluate our approach on the UNICORNv2 dataset, an EO and SAR dataset of civilian vehicles with heavily speckled images and severe class imbalance. EO prototype alignment improves SAR classification accuracy over frozen DINOv3, SAR only finetuning and unpaired distribution alignment baselines, and t-SNE visualizations provide qualitative evidence of clearer separation among classes in the trained SAR embedding space. These results suggest that optical vision foundation models, despite being trained on visible spectrum imagery, provide transferable information for SAR image classification, offering a practical method for using large scale pretrained vision foundation models across challenging sensing modalities.
Supervised synthetic aperture radar (SAR) automatic target recognition (ATR) methods rely on a closed-set assumption and struggle to recognize unseen target categories without labeled SAR samples. Zero-shot learning (ZSL) offers a promising solution by transferring knowledge from seen classes and external semantic prio...
Rui Zhu, Tian-Wen Zhang, Xiao-Ling Zhang et al.· IEEE Geoscience and Remote S...· 1 citation
Conventional deep-learning methods for synthetic aperture radar (SAR) target recognition predominantly rely on image-domain visual features, with insufficient integration of the physical principles governing SAR imaging. This limitation leads to unsatisfactory recognition accuracy and poor generalization performance un...
Aircraft often manifest as highly aspect/pose-sensitive, disjointed blobs made of pixels with fluctuating levels of brightness in classic grayscale SAR images, which makes the aircraft annotation task challenging even for human experts. As more and more high-resolution colorized SAR images acquired by the latest commer...
Yu Zhang, Zhe Geng, Lu-Jian Yao et al.· Remote Sensing· 0 citations
Automatic Target Recognition (ATR) in Synthetic Aperture Sonar (SAS) is a task largely dominated by deep neural networks (DNNs). Most SAS-ATR models use convolutional neural network (CNN) architectures whereas transformer-based architectures have had much less representation in the literature despite being state of the...
Paired synthetic aperture radar (SAR) and electro-optical (EO) imagery is increasingly available across sensors, resolutions, and geographic regions. Yet existing SAR-to-EO image translation (SET) methods are typically trained on a single, limited-scale dataset, producing models specialized to particular sensing condit...
Automatic classification of military aircraft in satellite imagery is a challenging problem with a high risk of error due to the limited pixel area of targets, variations in image resolution and illumination conditions, background complexity, and strong visual similarity among classes. In this study, a deep learning ap...
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Microsoft Research Blog· microsoft.comAug 11, 2026
Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.