Sequence-Aware Dataset Auditing for Leakage-Free Benchmarking of YOLO Detectors for Bottle Detection
Abstract
This work presents a reproducible YOLO-based pipeline for bottle detection in sandy environments, emphasizing dataset integrity, leakage-free evaluation, and deployment-oriented model selection. A one-class dataset of 1585 images and 3167 annotated bottles was audited to identify annotation-format defects and near-duplicate contamination between training and validation partitions. Sequence membership was reconstructed through perceptual-image similarity and used to assign complete image components to a sequence-aware train/validation split, eliminating the near-duplicate pairs found in the initial random partition. A controlled ablation holding model, seed, and corrected labels fixed showed that the random split reports 0.040 higher mAP@0.5:0.95 than the sequence-aware split (0.787 vs. 0.747), quantifying the leakage risk directly rather than only asserting it. Five YOLO configurations were then benchmarked under three independent seeds each; the observed mAP@0.5:0.95 differences among models (0.004–0.008) were small in absolute magnitude and, given only three seeds per model, are interpreted descriptively rather than as evidence of statistical equivalence or significance, so yolo11n_bottle was selected through a joint accuracy-parity, compactness, and exportability criterion (precision 0.982, recall 0.985, mAP@0.5 0.992, mAP@0.5:0.95 0.748), using approximately ten times fewer parameters than the largest configuration and producing a 5.2 MB checkpoint. ONNX export preserved detection geometry closely (100% count agreement, mean matched IoU ≥0.9998), without meeting strict metric-parity tolerances. A stratified sample of 108 frames from operational RealSense BAG footage was manually annotated by an independent reviewer and evaluated quantitatively: mAP@0.5 remained close to the internal validation figure (0.927 vs. 0.992), while mAP@0.5:0.95 fell substantially (0.483 vs. 0.747), revealing a localization gap between the curated benchmark and operational conditions that this manuscript reports transparently. Together, these results show that dataset auditing, sequence-aware partitioning, multiseed benchmarking, and manually annotated operational evidence are each necessary to interpret a detection benchmark built from continuous video acquisition, providing a traceable, reproducible workflow for selecting and evaluating compact visual-perception models for resource-constrained environmental applications.