It is shown that a better name repairs part of the gap, while the remaining failures call for an address that the tested natural-language rephrasings do not provide, and the model improves over few-shot methods that instead inject visual prompts at inference.
Abstract
Open-vocabulary segmentation labels arbitrary categories from a text query without per-class training, yet on remote sensing imagery it underperforms on categories it handles reliably elsewhere. We find that much of this gap traces to the text query rather than to the segmentation model. Because these models are not specialized for overhead imagery, the class name that serves as the query is often a weak address into the vision-language embedding space. We show that a better name repairs part of the gap, while the remaining failures call for an address that the tested natural-language rephrasings do not provide. We recover that address from a few examples through textual inversion on a frozen model, keeping inference text only. On a representative benchmark this raises the mean intersection over union on the affected categories from 3.9 to 39.4, and across eight remote sensing datasets it improves over few-shot methods that instead inject visual prompts at inference.
DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...
Mustafa Alawadi, M. Fateh· Jordanian Journal of Compute...· 0 citations
Open-vocabulary semantic segmentation (OVS) of remote sensing imagery is a challenging pixel-level task requiring strong generalization and adaptation to the spatial characteristics of remote sensing data. Although existing vision-language foundation models perform well in general domains, their image-level classificat...
Chang-Hao Zhao, Ling-Lin Zeng, Hai Liu· 0 citations
Open-vocabulary detectors pretrained on remote sensing imagery can be queried with free-form category names instead of a fixed taxonomy. In practice, they reach applications through fine-tuning on small, specialized target datasets, and standard evaluation reports only target-dataset accuracy, so any loss of the pretra...
Le-Fan Wang, Yuan-Rong He, Hui-Lin Xu et al.· IEEE Journal of Selected Top...· 0 citations
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches an...
Quan-Wei Liu, Tao Huang, Jia-Qi Yang et al.· 0 citations
This work formulates active learning for VG under the realistic setting where only raw images are available without accompanying text, and introduces Referred Region Ambiguity, a new acquisition function that measures whether the model's confidence collapses onto a single region or disperses across multiple candidates.
Junbeom Hong, Seonghoon Yu, Hyungsik Jung et al.· 0 citations
A novel framework that mines target knowledge using the strong reasoning capacity of Multimodal Large Language Models (MLLMs) and employs it to enhance FSS, which shows promising results and largely surpasses existing methods.
Yi-Jun Hu, Heng Fan, Li-Bo Zhang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.