Skip to content
Preprint

Class Geometry as Supervision for Sample-Efficient Open-World Detection

Aug 2026 · 0 citations · 22 references
Computer Science

TL;DR

Class-geometry supervision (CGS) is proposed, a general framework that constrains learned prototype or class-representation spaces to preserve visual or semantic class dissimilarities estimated from training data and suggests that relational class geometry is an effective supervisory signal for building calibrated and extensible open-world detectors under limited supervision.

Abstract

Open-world object detection requires models to recognize known categories, reject unfamiliar objects, and incorporate new classes over time. This is especially challenging in scarce-data settings such as biomedical and scientific imaging, where rare categories may have only a few annotated examples and fine-grained classes differ by subtle morphology. Prototype-based detectors are natural for this regime, but they typically learn class prototypes as independent anchors, ignoring relational structure among classes. We propose class-geometry supervision (CGS), a general framework that constrains learned prototype or class-representation spaces to preserve visual or semantic class dissimilarities estimated from training data. CGS introduces a dissimilarity-preserving objective that aligns pairwise distances among learned class representations with a target class-geometry matrix while retaining the standard task loss. We instantiate the same objective across prototype recognition, few-shot biomedical object detection, open-set detection, novel-class insertion, and OWOD adaptation on COCO. Experiments show that CGS improves sample efficiency in recognition and ova detection, substantially strengthens novel-class insertion, and improves unknown recall on COCO while retaining much of the known-class detection performance. Ablations show that meaningful visual geometry provides the most reliable gains, while random geometry can help novel separation but is less consistent for few-shot detection. These results suggest that relational class geometry is an effective supervisory signal for building calibrated and extensible open-world detectors under limited supervision.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Background-Free Objectness Learning for Class-Agnostic Detection

Object detectors are typically trained under closed-set supervision, where unlabeled regions are implicitly treated as background. Under incomplete annotations, this assumption introduces objectness bias: visually valid but unlabeled objects are used as negatives, tying objectness to the annotated taxonomy rather than...

Dania Batool, Liliana Lo Presti, M. La Cascia et al. · 0 citations
Preprint Aug 2026

Towards Sparsely Annotated Open-World Object Detection

Real-world object detection operates under ambiguous supervision, where unlabeled regions may correspond to missing annotations of known objects or genuinely unknown categories. These challenges have been addressed separately in Sparsely Annotated Object Detection (SAOD) and Open-World Object Detection (OWOD). In pract...

HeeJu Han, AJeong Kim, Jinsun Park · 0 citations
Preprint Aug 2026

WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes

Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is absent, large vision-language models usually address this task. We ask whether the same capability is available far more cheaply, from representations already learned by a...

K. Gupta, Ahmed Rafi Hasan, Md. Mahfuzur Rahman et al. · 0 citations
Sep 2026

Multimodal graph-based fusion via image descriptions for few-shot open-set recognition.

A multimodal Graph-based Fusion (MGF) framework that learns visually grounded semantic representations to enhance FSOR performance and achieves superior open-set recognition and competitive closed-set classification performance.

Xilang Huang, Seon-Han Choi · 0 citations
Open access Sep 2026

Protodetect: Prototype Compactness and Inference Calibration for Few-Shot Out-of-Distribution Detection

Detecting out-of-distribution (OOD) samples from limited labeled data is important for reliable recognition under semantic novelty. Existing few-shot approaches use synthetic or auxiliary unknowns, model only in-distribution (ID) data, or rely on large pretrained vision–language models. This paper presents Protodetect,...

Ze-Xia Huang, Hao-Yu Jiang, Jin-Song Hu et al. · 0 citations
Aug 2026

Self-supervised skeleton action recognition based on graph prototype learning

This work presents a novel self-supervised architecture centered on graph prototype learning that sets a new state-of-the-art on the ARMM dataset with an accuracy of 95.70%, substantiating the efficacy and transferability of prototype-guided self-supervised learning for skeleton-based action representation.

Zhijie Xu, Hong-Wei Chen, Xia Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.