Skip to content
Book Open access

OpenAqua: A Large-Scale Fine-Grained Dataset and Benchmark for Open Underwater Visual Perception

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 9522-9533 · 0 citations · 16 references

TL;DR

This work introduces OpenAqua, the first large-scale fine-grained dataset dedicated to open underwater visual tasks, and establishes a comprehensive benchmark suite that encompasses not only standard object detection and instance segmentation tasks but also pioneers an underwater open-vocabulary object detection benchmark.

Abstract

Monitoring aquatic biodiversity is vital for maintaining global ecological balance. While advancements in computer vision have revolutionized underwater perception, existing datasets are predominantly limited to coarse-grained categories or lack spatial localization annotations, severely constraining the applicability of models for fine-grained biological identification in real-world scenarios. To address this gap, we introduce OpenAqua, the first large-scale fine-grained dataset dedicated to open underwater visual tasks. OpenAqua is structured around a five-level biological taxonomic hierarchy, comprising 77,970 high-quality images covering 16,540 aquatic species, and providing 132,885 fine-grained bounding boxes and corresponding instance segmentation masks. Based on this dataset, we establish a comprehensive benchmark suite that encompasses not only standard object detection and instance segmentation tasks but also pioneers an underwater open-vocabulary object detection benchmark. Extensive experimental evaluations reveal a significant performance degradation in current models as they progress from coarse-grained to fine-grained recognition. These results highlight the substantial challenges associated with fine-grained semantic perception and domain adaptation in degraded underwater environments. We believe OpenAqua holds the potential to advance fine-grained underwater vision research, facilitate learning from long-tailed distributions, and enable more effective aquatic ecosystem monitoring. Our dataset is available at https://github.com/White-cat-ed/OpenAqua.

Read PDF

Similar papers

Review Open access Aug 2026

WetVeg-2mm: An Ultra-High-Resolution UAV Dataset for Riparian Vegetation Semantic Segmentation

Fine-grained mapping of riparian vegetation is important for ecological monitoring, invasive species control, and ecosystem restoration. However, riparian plant communities often exhibit fragmented patches, broad transition zones and high visual similarity among classes, making stable species-level segmentation difficult from conventional satellite imagery or lower-resolution UAV imagery. To address this gap, we present WetVeg-2mm, an ultra-high-resolution UAV dataset for fine-grained riparian vegetation semantic segmentation. Built from UAV surveys over a representative riparian section of the Jiuzhou River in Guangxi, China, the dataset provides 2054 image chips (1024 × 1024) with pixel-level annotations at 2 mm ground sampling distance. It contains 17 semantic classes in total, including 14 representative wetland plant classes, such as Colocasia, Eichhornia and Phragmites, together with water, bareland and background. Five baseline models, namely U-Net, Attention U-Net, DeepLabV3+, PSPNet and SegFormer, were evaluated using per-class IoU, mIoU, mDice, PA, Precision and Recall. Across all evaluated baseline settings, SegFormer with ImageNet pretraining achieved the best overall performance, with 76.23% mIoU, 86.01% mDice, 85.17% PA, 87.30% Recall and 85.42% Precision on the test set. Overall, WetVeg-2mm provides a reproducible and challenging benchmark for fine-grained riparian vegetation semantic segmentation.

Guiqi Liu, Runqiao Zhang, Huapeng Qin · 0 citations
Preprint Aug 2026

OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization

Fine-grained Cross-View Localization (CVL) estimates the precise position and orientation of a ground-level image by aligning it with geo-referenced aerial imagery, offering a scalable alternative to Global Navigation Satellite Systems (GNSS) in challenging urban environments. Existing datasets rely on data collected with high-end sensor suites, which inherently limit image diversity and scalability. While in-the-wild images are abundant, their noisy geo-tags make them unsuitable for reliable evaluation. To bridge this gap, we introduce OpenCVL, a large-scale, diverse, and open dataset containing 617,388 ground-aerial image pairs spanning 41 cities across four European countries. All images are sourced from permissive platforms, ensuring long-term accessibility and supporting open and reproducible research. The training set combines images captured with high-end sensors with diverse in-the-wild imagery. We further develop a data curation framework that filters and corrects pose annotations to construct reliable in-the-wild evaluation data. In addition, OpenCVL includes dedicated cross-area and snowy test sets to assess generalization and robustness. Experiments with a state-of-the-art CVL model on OpenCVL show that incorporating noisy in-the-wild data consistently improves performance on clean test sets, suggesting a promising direction for scaling CVL with diverse real-world imagery.

Zimin Xia, Mubariz Zaffar, Junfan Fu et al. · 0 citations
Review Jul 2026

EcoVision: AI-Powered Drone Imaging for Salt Marsh Vegetation Monitoring and Dominance Mapping

High-resolution RGB imagery acquired from low-altitude UAV surveys was processed through a modular pipeline incorporating transformer-based semantic segmentation, connected-component vegetation extraction, fine-grained species classification using a ConvNeXt architecture, and grid-based dominance scoring at 2x2m resolution. The framework targeted two ecologically significant halophytic grasses, Spartina maritima and Puccinellia maritima, and was trained using a curated and manually annotated UAV imagery, along with biodiversity imagery sourced from publicly accessible datasets. In order to identify these plants from the imagery, our segmentation yielded reliable species masks (mean IoU = 0.56; pixel-level accuracy = 0.96), while object-level classification achieved very good discrimination (F1 = 0.99). Dominance estimates closely matched quadrat-based field surveys, with mean absolute differences below 8%, preserving fine-scale spatial structure under realistic survey conditions. The developed system, named EcoVision, establishes a practical foundation for scalable, high-resolution salt marsh monitoring, demonstrating how AI-driven workflows can translate pixel-level predictions into ecologically interpretable metrics.

I. Onyenonachi, Peter J. Lawerance, Nadia Kanwal · 0 citations
Review Open access Aug 2026

WIO-ReefFish: A High-Resolution Dataset for Taxon-Aware Coral Reef Fish Detection in the Western Indian Ocean

Coral reef fish assemblages are widely used as indicators of ecosystem condition, yet manual annotation of underwater video remains a major bottleneck for scalable biodiversity monitoring. Despite rapid progress in automated detection, ecologically realistic and publicly available datasets remain scarce, particularly for the Western Indian Ocean. Here, we present WIO-ReefFish, a reef fish detection dataset derived from diver-operated line-intercept transects and designed for ecological monitoring under natural survey conditions. WIO-ReefFish comprises 1,000 ultra-high-definition images (3840 × 2160 pixels) and 6,768 exhaustive bounding-box annotations spanning 24 taxonomic categories, thereby preserving full-frame assemblage structure in complex reef scenes. We also establish a standardized benchmark across nine object detection models under two complementary protocols: class-aware detection and class-agnostic fish localization. Detection performance was consistently higher under the class-agnostic protocol. The best-performing model (RT-DETR) improved from 0.48 mAP50 in the class-aware setting to 0.70 mAP50 when taxonomic constraints were removed, indicating that taxonomic discrimination remains substantially more challenging than fish localisation in reef imagery. Spatially independent evaluation revealed a pronounced generalisation gap, particularly for taxonomic detection, whereas class-agnostic fish localisation remained substantially more robust across transects and countries. Together, these results establish WIO-ReefFish as a realistic benchmark for automated reef fish detection and provide a foundation for more robust computer-vision tools in coral reef biodiversity monitoring. The WIO-ReefFish dataset and associated benchmarking resources are publicly available.

J. Gerard, Luca Branger, F. Huyghe et al. · 0 citations
Jul 2026

LC-YOLO: long-tailed infrared UAV detection based on local contrast

Infrared UAV detection and fine-grained recognition are pivotal for low-altitude security but suffer from two primary issues: the lack of texture in infrared imagery, which hinders fine-grained classification, and extreme long-tailed data distributions, which leads to poor performance on rare classes. We propose LC-YOLO, a real-time framework integrating local contrast enhancement and category balancing. To resolve texture deficiency, we design a Multi-scale Local Contrast Module (MLCM) that utilizes dilated convolutions to mimic the Human Visual System, significantly enhancing rotor edge features for better fine-grained discrimination. To mitigate data imbalance, we introduce a Category-Specific Mosaic (CS-Mosaic) strategy that enforces tail-class oversampling during data loading, preventing model overfitting to head classes at the source. Experiments on a multi-source heterogeneous dataset demonstrate that LC-YOLO substantially improves overall mAP and tail-class recall (e.g., single-rotors and fixed-wings) while maintaining real-time efficiency.

Sen Song, Weida Zhan, Xuhao Liu et al. · 0 citations