Skip to content

Multimodal Semantic-Probabilistic Objectness for Open World Object Detection

Jul 2026 · arXiv.org · Vol abs/2607.23981 · 0 citations · 38 references
Computer Science

TL;DR

MSPO is proposed, a lightweight semantic calibration framework that augments PROB with task-aware known-category language priors while preserving its detector architecture and incremental learning protocol and demonstrates that known-category language semantics provide an effective calibration signal for probabilistic objectness under the standard OWOD setting.

Abstract

Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and incrementally learn newly annotated classes. PROB improves unknown discovery by modeling class-agnostic probabilistic objectness in the decoder-query space. However, visual objectness alone cannot determine whether an object-like query corresponds to a hard known instance, an unseen-category object, or background clutter, resulting in an ambiguous known-unknown decision boundary. We propose MSPO, a lightweight semantic calibration framework that augments PROB with task-aware known-category language priors while preserving its detector architecture and incremental learning protocol. For each currently known category, MSPO constructs an extended text description covering category attributes, visual appearance, typical scenes, and functional usage, and encodes it using a frozen CLIP text encoder. Decoder query features are projected into the same semantic space to estimate their support from the current known-category semantics. This semantic evidence is fused with PROB's visual objectness to calibrate known and unknown predictions without turning OWOD into open-vocabulary classification. Importantly, MSPO never uses future-category names, and all unseen categories remain unnamed during evaluation. Experiments on M-OWODB and S-OWODB show that MSPO improves the strong PROB baseline on the main aggregate metrics while retaining competitive unknown recall. It also improves early unknown-confusion metrics and raises PASCAL VOC final mAP by up to 2.7 points. These results demonstrate that known-category language semantics provide an effective calibration signal for probabilistic objectness under the standard OWOD setting.

View source

Similar papers

Preprint Aug 2026

Towards Sparsely Annotated Open-World Object Detection

Real-world object detection operates under ambiguous supervision, where unlabeled regions may correspond to missing annotations of known objects or genuinely unknown categories. These challenges have been addressed separately in Sparsely Annotated Object Detection (SAOD) and Open-World Object Detection (OWOD). In pract...

HeeJu Han, AJeong Kim, Jinsun Park · 0 citations
#artificial intelligence Preprint Aug 2026

Background-Free Objectness Learning for Class-Agnostic Detection

Object detectors are typically trained under closed-set supervision, where unlabeled regions are implicitly treated as background. Under incomplete annotations, this assumption introduces objectness bias: visually valid but unlabeled objects are used as negatives, tying objectness to the annotated taxonomy rather than...

Dania Batool, Liliana Lo Presti, M. La Cascia et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Open-World Semantic Segmentation with Sensitivity Modeling

This work addresses open-world semantic segmentation, the joint task of segmenting known classes while detecting and grouping novel or anomalous content without additional supervision, by extending a dual-decoder baseline with a third, complementary decoder within a unified encoder-decoder design.

Anastasios Romanos Varvarigos, Nikos Giakoumoglou, Tania Stathaki · 0 citations
Preprint Aug 2026

Class Geometry as Supervision for Sample-Efficient Open-World Detection

Class-geometry supervision (CGS) is proposed, a general framework that constrains learned prototype or class-representation spaces to preserve visual or semantic class dissimilarities estimated from training data and suggests that relational class geometry is an effective supervisory signal for building calibrated and...

A. Rao, Zhou Chen, Revanth Reddy Palem et al. · 0 citations
Preprint Aug 2026

CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection

Open World Object Detection (OWOD) built on multimodal foundation models often suffers from semantic ambiguity caused by unidirectional text-to-vision matching, while rigid outlier penalties may over-suppress unknown objects near known-class decision boundaries. We propose CODE (Cross-Modal Calibration and Dynamic Supp...

Hao Xu, Zhao-Ning Shi, Heng-Yu Jin et al. · 0 citations
Aug 2026

PointPDF V2: A Unified Framework for Continual Open-World 3D Semantic Segmentation.

This work proposes PointPDF V2, a unified framework that integrates open-set recognition (OSS) and incremental learning (IL) into a cohesive pipeline and introduces a more challenging continual OWSS protocol in 3D, where models must simultaneously preserve the known-class performance, acquire new knowledge, and still i...

Jinfeng Xu, Xianzhi Li, Yixue Hao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.