Considering the difficulty of learning spatially and semantically aware prompt injection, the Hierarchical Prompt Injector is proposed, which enables spatially adaptive prompt injection in foundation models and auxiliary supervision to align hierarchical prompts with their corresponding object regions is introduced.
Abstract
Domain Generalized Semantic Segmentation (DGSS) is a challenging task, as vision models often rely on low-level appearance cues that change across domains. In contrast, structural attributes exhibit cross-domain stability, motivating the use of structural priors for DGSS. Existing methods use prompt learning to transfer such priors into DGSS models, but typically encode each class as a single holistic prompt. Moreover, these methods apply prompts uniformly to all pixels, offering no mechanism to adapt when only a subset of object regions is visible due to viewpoint changes, occlusion, and environmental variation. We address this with \textbf{Spatial Hierarchical Prompts (SHP)} that enrich each class with region-level geometric anchors capturing structural appearance from distinct viewing angles, ensuring complementary coverage under arbitrary viewpoints. Additionally, we propose the \textbf{Hierarchical Prompt Injector (HPI)}, which enables spatially adaptive prompt injection in foundation models. HPI spatially grounds prompts by modeling their semantic relevance and spatial influence with visual features. Considering the difficulty of learning spatially and semantically aware prompt injection, we further introduce auxiliary supervision to align hierarchical prompts with their corresponding object regions. We achieve 70.62\% and 72.74\% mIoU on synthetic-to-real and real-to-real benchmarks, respectively. Code and checkpoints are released at https://github.com/MosukFate/HPI
Few-shot segmentation (FSS) remains a significant challenge due to the scarcity of annotated data and the need for precise object localization across novel classes. Existing approaches often rely on single-backbone architectures and coarse priors, which struggle to capture detailed semantics and precise spatial alignme...
Guo-Hua Geng, Xiao-Feng Wang· International Conference on...· 0 citations
The Language-and-Source-Anchored Alignment (LASA) framework is proposed, which comprises three synergistic components: Text-and-Source-Guided Style Transfer (TSGST), Domain-Aware Query Adapter (DAQA), and Domain-Aware Decoder Optimizer (DADO).
Jin-Hong Zhu, Wei-Qi Yan, Sheng-Chuan Zhang et al.· 0 citations
A complementary prototype representation framework is proposed, employing three modules to collaboratively improve pseudo-label quality and improves the discriminative ability of confused categories by generating semantically similar sub-category negative samples.
Open-world entity segmentation aims to predict masks for arbitrary objects across domains and at multiple granularities, from parts to whole objects. In this setting, SAM sets a strong standard: trained on SA-1B, comprising 11M images and over 1B carefully annotated masks, it achieves remarkable zero-shot performance....
Christoph Hümmer, Joachim Sicking, Fabian Hüger et al.· 0 citations
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absenc...
J. del Pino, Salvador Rodríguez, Alejandro Garabito et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.