We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as target entities. ENEAS works two ways from a single method: precise tracking and high-quality segmentation of a unique instance, and open-concept discovery of every instance a text query names, resolved by a semantic verification layer. For tracking, we extend the geometrically robust SeC architecture, previously limited to point interactions, with a text-prompting adapter and leverage its temporal memory, so that the target is held through disappearance without drifting to distractors and kept whole even when it fills the entire view. For discovery, the verification layer combines high-speed visual embedding matching with conditional VLM refinement, invoking semantic reasoning only for ambiguous candidates, which filters out the ontological errors that visual-only models cannot distinguish while keeping latency low. Designed with 3D reconstruction in mind, where a single misclassified distractor corrupts the asset, ENEAS unlocks high-quality semantic tracking and segmentation of video, of broad libraries, and of collections of temporally or spatially unordered data, together with the discrimination to tell true instances from their doppelgangers: things that look alike but are not the same. The code and models are available at https://github.com/speridlabs/eneas
Vision-language models show promise in zero-shot semantic segmentation, but a key challenge is the disconnect between text and visual features. While text embeddings can roughly localize unseen objects, they often lack the fine-grained detail necessary for accurate segmentation, leading to oversegmentation or undersegm...
Jia-Xiang Fang, Shi-Qiang Ma, Jing Wang et al.· Neural Networks· 0 citations
Open-world entity segmentation aims to predict masks for arbitrary objects across domains and at multiple granularities, from parts to whole objects. In this setting, SAM sets a strong standard: trained on SA-1B, comprising 11M images and over 1B carefully annotated masks, it achieves remarkable zero-shot performance....
Christoph Hümmer, Joachim Sicking, Fabian Hüger et al.· 0 citations
A novel framework that mines target knowledge using the strong reasoning capacity of Multimodal Large Language Models (MLLMs) and employs it to enhance FSS, which shows promising results and largely surpasses existing methods.
Considering the difficulty of learning spatially and semantically aware prompt injection, the Hierarchical Prompt Injector is proposed, which enables spatially adaptive prompt injection in foundation models and auxiliary supervision to align hierarchical prompts with their corresponding object regions is introduced.
Xin Lin, Ruo-Yu Guo, Jia-Qi Guo et al.· 0 citations
FoRIS consists of three key stages: Foreground Purification, Foreground Localization, and Foreground Consolidation, which progressively suppress background distractions, localize discriminative target regions, and recover complete foreground structures through semantic aggregation.
Ming Hu, Jian-Fu Yin, Ming-Yu Dou et al.· 0 citations
Personalized segmentation and personalized retrieval both aim to identify the same physical object across different images. While the former localizes the object within a target image, the latter retrieves images where it appears. Despite this shared instance-level objective, the two tasks have largely evolved separate...
Gabriele Trivigno, Marcos Alfaro, C. Cuttano et al.· 0 citations
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.