This work introduces a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object, and shows that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable.
Abstract
While object detection has advanced through improved architectures and open-vocabulary models, we provide strong evidence that benchmark quality is limited by annotation incompleteness. Across four widely used datasets (COCO, Pascal VOC, Cityscapes, KITTI), re-annotation reveals substantial increases in annotated objects (e.g., up to +60% on KITTI and +40% on COCO), driven primarily by previously unlabeled small, occluded, or densely packed instances. While some differences arise from dataset-specific annotation conventions, we consistently find that missing annotations are the main source of label errors across all datasets. To achieve high data quality, we introduce a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object. The resulting annotations improve coverage and align well with human calibration. We show that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable. We introduce two large-scale benchmarks: (i) an uncertainty-aware object detection benchmark, and (ii) a label error detection benchmark grounded in real label errors. We show that current detectors are strongly depended on annotation quality and are misaligned with human perception. Current label error detection methods, which have been shown to perform well on synthetic noise, struggle to achieve high recall and precision on real label errors. Our results highlight the need for future object detection benchmarks to move beyond deterministic annotations toward high-recall, uncertainty-aware evaluation that maximizes valid instances and better reflects real-world ambiguity.
A five-axis taxonomy (modality, mechanism, prompting, supervision level, and generalization setting) is introduced to audit the literature across application domains, including microscopy, remote sensing, crowd counting, and agriculture, and formalizes prevailing challenges into six structural contradictions.
Joana Konadu Owusu, S. Sheshappanavar· 0 citations
Precise image annotation is essential for training advanced computer vision models. However, making high-accuracy, pixel-level image annotation efficient is never easy: manual labeling is highly accurate but time-consuming, while automatic labeling that using state-of-the-art models is much faster but often yields unsa...
Sheng-Qing Xia, Jia-Xin Du, Chun-Yi Peng· Proceedings of the 4th Inter...· 0 citations
Point-based weakly supervised strategies can reduce annotation costs but suffer from the problems of semantic sparsity and boundary ambiguity. This paper proposes an uncertainty-aware weakly supervised building change detection method with point annotations. First, the method leverages the proposal generation capabilit...
Yongqing Wang, Er-Zhu Li, Lihua Du et al.· Photogrammetric Engineering...· 0 citations
In semi-supervised object detection (SSOD), due to the limited availability of labeled data, the quality and quantity of pseudo labels generated from unlabeled images are crucial for model training. Our study reveals that in the early stages of training, the number of usable pseudo labels is very low, which hampers the...
Xi Yang, Peng-Hui Li, Nan-Nan Wang· IEEE Transactions on Image P...· 0 citations
A novel neural network architecture that accepts natural language-labeled keypoint graphs as prompts and predicts the image coordinates of graph nodes and achieves strong performance on the public MP-100 benchmark while offering greater flexibility in representing and localizing complex objects.
Patrick Tirler, J. Piater· Machine Vision and Applicati...· 0 citations
This work formulates active learning for VG under the realistic setting where only raw images are available without accompanying text, and introduces Referred Region Ambiguity, a new acquisition function that measures whether the model's confidence collapses onto a single region or disperses across multiple candidates.
Junbeom Hong, Seonghoon Yu, Hyungsik Jung et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.