Skip to content
Preprint

Object Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty

Sep 2026 · 0 citations
Computer Science

TL;DR

This work introduces a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object, and shows that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable.

Abstract

While object detection has advanced through improved architectures and open-vocabulary models, we provide strong evidence that benchmark quality is limited by annotation incompleteness. Across four widely used datasets (COCO, Pascal VOC, Cityscapes, KITTI), re-annotation reveals substantial increases in annotated objects (e.g., up to +60% on KITTI and +40% on COCO), driven primarily by previously unlabeled small, occluded, or densely packed instances. While some differences arise from dataset-specific annotation conventions, we consistently find that missing annotations are the main source of label errors across all datasets. To achieve high data quality, we introduce a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object. The resulting annotations improve coverage and align well with human calibration. We show that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable. We introduce two large-scale benchmarks: (i) an uncertainty-aware object detection benchmark, and (ii) a label error detection benchmark grounded in real label errors. We show that current detectors are strongly depended on annotation quality and are misaligned with human perception. Current label error detection methods, which have been shown to perform well on synthetic noise, struggle to achieve high recall and precision on real label errors. Our results highlight the need for future object detection benchmarks to move beyond deterministic annotations toward high-recall, uncertainty-aware evaluation that maximizes valid instances and better reflects real-world ambiguity.

View source

Similar papers

Review Aug 2026

Object Counting Across Modalities: Taxonomies, Benchmarks, Applications, and Open Challenges

A five-axis taxonomy (modality, mechanism, prompting, supervision level, and generalization setting) is introduced to audit the literature across application domains, including microscopy, remote sensing, crowd counting, and agriculture, and formalizes prevailing challenges into six structural contradictions.

Joana Konadu Owusu, S. Sheshappanavar · 0 citations
Book Open access Oct 2026

FollowMe: Automating Precise Image Annotation Across Various Environments

Precise image annotation is essential for training advanced computer vision models. However, making high-accuracy, pixel-level image annotation efficient is never easy: manual labeling is highly accurate but time-consuming, while automatic labeling that using state-of-the-art models is much faster but often yields unsa...

Sheng-Qing Xia, Jia-Xin Du, Chun-Yi Peng · 0 citations
2026

Uncertainty-Aware Weakly Supervised Building Change Detection with Point Annotations

Point-based weakly supervised strategies can reduce annotation costs but suffer from the problems of semantic sparsity and boundary ambiguity. This paper proposes an uncertainty-aware weakly supervised building change detection method with point annotations. First, the method leverages the proposal generation capabilit...

Yongqing Wang, Er-Zhu Li, Lihua Du et al. · 0 citations
Sep 2026

From Generation to Optimization: Improving Pseudo Labels for Semi-Supervised Object Detection

In semi-supervised object detection (SSOD), due to the limited availability of labeled data, the quality and quantity of pseudo labels generated from unlabeled images are crucial for model training. Our study reveals that in the early stages of training, the number of usable pseudo labels is very low, which hampers the...

Xi Yang, Peng-Hui Li, Nan-Nan Wang · 0 citations
Open access Aug 2026

Natural language-labeled keypoint graphs for industrial object localization

A novel neural network architecture that accepts natural language-labeled keypoint graphs as prompts and predicts the image coordinates of graph nodes and achieves strong performance on the public MP-100 benchmark while offering greater flexibility in representing and localizing complex objects.

Patrick Tirler, J. Piater · 0 citations
#artificial intelligence Preprint Aug 2026

Cost-efficient Active Learning for Referring Image Segmentation and Grounding

This work formulates active learning for VG under the realistic setting where only raw images are available without accompanying text, and introduces Referred Region Ambiguity, a new acquisition function that measures whether the model's confidence collapses onto a single region or disperses across multiple candidates.

Junbeom Hong, Seonghoon Yu, Hyungsik Jung et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.