A novel mechanism to automatically identify which of these point-labels are suitable, and which are actively harmful, when used for propagation is introduced, paving the way for scalable ecological analysis.
Abstract
The health of marine ecosystems is a critical indicator of global environmental change, yet the physical constraints of underwater observation and the intrinsic challenges of processing marine imagery severely limit the scalability of systematic monitoring. While recent visual foundation models such as the Segment Anything Model (SAM) series show great promise, they still struggle with the fine-grained recognition required in these complex scenarios and still require expert supervision. Our work addresses this gap by bridging state-of-the-art foundation models with existing sparse supervision. Because historical benthic surveys are typically annotated with only a few sparse expert points per image, we utilize these legacy point-labels as visual prompts for SAM2. Our primary contribution is a novel mechanism to automatically identify which of these points are suitable, and which are actively harmful, when used for propagation. By filtering out unreliable points, we extract high-quality pseudo-ground-truth masks capable of training more accurate, fine-grained semantic segmentation models. We demonstrate the effectiveness of our approach on public benthic data and introduce a new, challenging benchmark featuring real-world sparse expert annotations, paving the way for scalable ecological analysis.
Abstract. Semantic segmentation of remote sensing imagery (RSI) is essential for urban mapping, land-use monitoring, and many other domains. However, pixel-level annotation is expensive, making weakly supervised semantic segmentation (WSSS) that relies on image-level labels an attractive alternative. Pre-trained models provide strong priors from large-scale learned representations, making them beneficial for WSSS. However, when kept frozen, they often produce sparse and misaligned class activation maps (CAMs) due to domain gaps and static inference. We propose a lightweight and efficient framework that integrates CLIP and DINO foundation models to address three challenges: (i) semantic misalignment between generic text prompts and RSI-specific visuals; (ii) static CAM quality; and (iii) incomplete object coverage. Our design includes: (1) a Textual Prototype-Aware Enrichment (TPE) module that builds an RS-specific knowledge base using large language model (LLM)-generated descriptions to enrich text prompts; (2) a Unified Semantic Relation Mining (USR) module that fuses learnable adapter features with CLIP attention and DINO affinity for online CAM refinement; and (3) a Visual Prototype-Aware Enrichment (VPE) module, which maintains momentum visual prototypes to complete regions and sharpen boundaries. By freezing the CLIP and DINO backbones and optimizing only lightweight adapter and decoder modules, the proposed framework reduces the number of trainable parameters while achieving competitive performance. Experimental on iSAID and ISPRS Potsdam datasets demonstrate the effectiveness of the proposed framework, achieving 38.01% mIoU on iSAID dataset and 47.01% mIoU with 66.89% overall accuracy on Potsdam dataset.
Xin Li, Nicola Genzano, M. Gianinetto et al.· ISPRS Annals of the Photogra...· 0 citations
Experimental results demonstrate that the adapted SAM2 model achieves stable segmentation under moderate environmental variability, while degrading under severe visibility loss, consistent across model scales and input resolutions.
Bindusara Nagathihalli Lokesh, Laura Camila Duran Vergara, Hans-Gerd Maas et al.· The International Archives o...· 1 citation
The first systematic zero-shot evaluation of SAM 2 for aerial building segmentation is presented, establishing SAM 2 as a viable tool for rapid building mapping while highlighting where domain adaptation remains necessary.
Bingning Xiong, Mingyu Ou· Journal of image processing...· 0 citations
This work introduces OpenAqua, the first large-scale fine-grained dataset dedicated to open underwater visual tasks, and establishes a comprehensive benchmark suite that encompasses not only standard object detection and instance segmentation tasks but also pioneers an underwater open-vocabulary object detection benchmark.
Linxuan Luo, Pan Mu, Cong Bai· Proceedings of the 32nd ACM...· 0 citations
This dual-path design, tailored to the geometric characteristics of buildings, is the first attempt to fully exploit SAM’s complementary capabilities in a unified training-free pipeline, exhibiting excellent accuracy, robustness, and cross-dataset adaptability, and providing valuable insights for practical remote sensing applications.
Junming Chen, Bing Liu, Weiqi Lian et al.· IEEE Journal of Selected Top...· 0 citations
The proposed GeoRGMAE, a geospatially guided masked autoencoder pretraining strategy for building segmentation, introduces three masking strategies that prioritize semantically relevant building regions under the varying urban densities and suggests that incorporating geospatial priors into masked image modelling (MIM) can improve representation learning for downstream building segmentation tasks.
Tuğba Eraslanoğlu, G. Mutreja, Martin Kada et al.· The International Archives o...· 0 citations
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.