Jul 2026
DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
DINOde is proposed, an ODE-based framework that continuously aligns CLIP text embeddings with the DINO visual manifold through a continuous ODE trajectory, and achieves state-of-the-art performance across multiple OVSS benchmarks.
Sung-Hoon Yoon, Hoyong Kwon, Chang-Hwan Oh et al.
· arXiv.org · 0 citations