Object Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty
This work introduces a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object, and shows that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable.