Self-visual prompting (SVP), a training-free CAM generation paradigm that employs a recursive self-bootstrapping strategy to shift optimization from model parameters to the input context, and progressively refines the quality of CAMs.
Considering the difficulty of learning spatially and semantically aware prompt injection, the Hierarchical Prompt Injector is proposed, which enables spatially adaptive prompt injection in foundation models and auxiliary supervision to align hierarchical prompts with their corresponding object regions is introduced.
Xin Lin, Ruo-Yu Guo, Jia-Qi Guo et al.· 0 citations
Weakly Supervised Semantic Segmentation (WSSS) learns pixel-level predictions from image-level tags. Recent work focuses on improving coarse CAMs extracted from large vision-language models (commonly CLIP), but does little to improve their accuracy along segment boundaries. That job is instead delegated to a post-proce...
Recent advances in vision foundation models (VFMs) have shown remarkable capabilities across diverse unimodal visual tasks. However, adapting VFMs to referring image segmentation (RIS) typically necessitates precise vision-language alignment via full fine-tuning, incurring substantial computational overhead and risking...
Xiaoqiang Lu, Li-Cheng Jiao, Ling-Ling Li et al.· 0 citations
Open-vocabulary semantic segmentation (OVSS) enables pixel-level prediction over arbitrary text-specified vocabularies and has shown strong generalization on common benchmarks. However, OVSS performance often degrades in specialized domains such as medical imaging, remote sensing, and industrial inspection, where dense...
Hyun-Kurl Jang, Jihun Kim, Kuk-Jin Yoon· 0 citations
Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...
Trong-Tai Dam Vu, Vinh-Tiep Nguyen· International Conference on...· 0 citations
Anatomically plausible segmentation remains challenging because of low contrast, ambiguous boundaries, and modality-specific artifacts. Interactive segmentation has emerged as a promising strategy to guide feature extraction and improve localization, particularly in structurally ambiguous regions. However, existing met...
Mosharof Hossain, Md. Rabiul Islam, Limon Halder et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.