Similarity-Shift Refinement (SSR), a training-free post-hoc method for improving object-centric masks with a frozen self-supervised Vision Transformer, provides a simple and transferable signal for training-free object-centric mask refinement.
Abstract
Object-centric models often produce fragmented masks, boundary leakage, and incorrect region merging. We introduce Similarity-Shift Refinement (SSR), a training-free post-hoc method for improving object-centric masks with a frozen self-supervised Vision Transformer. SSR measures changes in pairwise patch similarity before and after self-attention value aggregation, retains positively strengthened relations, and constructs a sparse affinity graph. This graph propagates the initial soft slot assignments in a single refinement step, without retraining or modifying either model. Across natural-image, synthetic-video, and real-world-video benchmarks, SSR improves all-pixel Adjusted Rand Index in all 24 evaluated model-dataset combinations, with an average gain of 8.5 percentage points. Ablations show that value-space similarity shifts outperform query- and key-space variants as well as static Transformer affinities. However, texture-dense scenes may cause visually similar regions to be over-grouped. Overall, SSR provides a simple and transferable signal for training-free object-centric mask refinement.
Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...
Trong-Tai Dam Vu, Vinh-Tiep Nguyen· International Conference on...· 0 citations
UnfoldCRF treats refinement as inference in a conditional random field over pixel labels and latent region variables, and zeroing individual messages shows where the corrections come from: the pairwise messages mostly fix boundaries, the region messages mostly fix non-boundary errors.
Chun-Ming He, Rihan Zhang, Lei Xu et al.· 0 citations
PredErase, a training-free inference procedure on frozen FLUX.2 and I-JEPA, improves the native FLUX.2 backbone and is supported is training-free object-and-effect editing of frozen Fill, not replacement of paired-data erasers.
Wai-Kit Xiu, Qiang Lu, Jun-Biao Chen et al.· 0 citations
Results suggest that decoupled matching improves robustness to appearance-driven confusion in indoor point-cloud segmentation and propose D2M-Net, a decoupled dual-matching network that separates backbone features into geometry-oriented and semantic-oriented subspaces before prototype comparison.
Han-Bin Fang, Cheng-Long Peng, Xue-Yong Xiang et al.· The Visual Computer· 0 citations
This paper proposes a zero-shot framework for constrained latent inpainting with a frozen pretrained Stable Diffusion model, requiring no task-specific training or model fine-tuning and demonstrates effective object removal and context-consistent replacement content.
Arman Taghizadeh, U. Krumnack, Kai-Uwe Kühnberger· 0 citations
This work proposes a region-based feature enhancement framework built upon a Topology-aware Segment Graph (TSG) that achieves superior color accuracy and temporal stability compared to state-of-the-art frameworks, particularly in scenarios involving complex character motion and topological variation.
Bin Huang, Haoran Mo, Chengying Gao· IEEE Transactions on Visuali...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.