Skip to content
Dataset Open access

A Dataset for Fish Segmentation and Tracking in Underwater Videos

Aug 2026 · Scientific Data · Vol 13 · 0 citations · 40 references
Medicine

TL;DR

An open-source, browser-based annotation tool integrating the Segment Anything Model (SAM2) and CUTIE for efficient semi-automatic segmentation and tracking and facilitates high-quality annotations without specialized hardware, improving accessibility and reproducibility within the marine imaging community.

Abstract

The automatic monitoring of fish in underwater imagery plays a key role in marine ecology, fisheries management, and environmental monitoring, yet progress is limited by the lack of large, high-quality fish-focused datasets. We present a new dataset of underwater videos of fish in natural habitats, annotated for pixel-level segmentation and multi-object tracking. The data was collected in the Balearic Sea, the western Mediterranean region surrounding the island of Mallorca (Spain), across diverse marine environments to capture variations in species, lighting, turbidity, and background complexity. Each video frame has been carefully annotated to ensure spatial and temporal consistency, yielding a challenging and comprehensive resource for developing and benchmarking underwater vision algorithms. To illustrate the dataset’s utility, we provide baseline tracking results obtained with Deep OC-SORT, which highlight both the dataset’s challenging nature and its potential for future method evaluation. In addition, we release an open-source, browser-based annotation tool integrating the Segment Anything Model (SAM2) and CUTIE for efficient semi-automatic segmentation and tracking. This tool facilitates high-quality annotations without specialized hardware, improving accessibility and reproducibility within the marine imaging community.

Read PDF

Similar papers

Review Jul 2026

Under Water Animal Detection using CNN And Large Language Processing

ABSTRACT  Underwater computer vision plays a vital role in ocean research, enabling autonomous navigation, infrastructure inspections, and marine life monitoring. However, the underwater environment presents unique challenges, including color distortion, limited visibility, and dynamic light conditions, which hinder the performance of traditional image processing methods. Recent advancements in deep learning (DL) have demonstrated remarkable success in overcoming these challenges by enabling robust feature extraction, image enhancement, and object recognition. This review provides a comprehensive analysis of cutting-edge deep learning architectures designed for underwater object detection, segmentation, and tracking. State- of-the-art (SOTA) models, including AGW-YOLOv8, Feature-Adaptive FPN, and Dual-SAM, have shown substantial improvements in addressing occlusions, camouflaging, and small underwater object detection. For tracking tasks, transformer-based models like SiamFCA and FishTrack leverage hierarchical attention mechanisms and convolutional neural networks (CNNs) to achieve high accuracy and robustness in dynamic underwater environments. Beyond optical imaging, this review explores alternative modalities such as sonar, hyperspectral imaging, and event-based vision, which provide complementary data to enhance underwater vision systems. These approaches improve performance under challenging conditions, enabling richer and more informative scene interpretation. Promising future directions are also discussed, emphasizing the  need for domain adaptation techniques to improve generalizability, lightweight architectures for real-time performance, and multi-modal data fusion to enhance interpretability and robustness. By critically evaluating current methodologies and highlighting gaps, this review provides insights for advancing underwater computer vision systems to support ocean exploration, ecological conservation, and disaster management. INDEX TERMS  Underwater computer vision, deep learning, underwater robotics, ocean research, underwater image enhancement, object tracking, object detection.

CH Vasundara, G. M. Krishna · 0 citations
Conference Aug 2026

Collaborative video object segmentation for underwater filter-net inspection

Underwater filter-nets play a critical role in aquaculture and marine engineering, where reliable condition monitoring is essential for ensuring operational safety. Although video-based filter-net segmentation enables automated inspection and early fault detection, its performance is significantly hindered by underwater imaging challenges, including low illumination, scattering-induced visibility degradation, and pronounced spatiotemporal appearance variability. These challenges often cause conventional segmentation approaches to exhibit mask drift and error accumulation, thereby compromising stable long-term tracking. To address these challenges, we propose an enhanced SAM2-based segmentation framework incorporating two collaborative temporal-consistency mechanisms that combines mask-weakening and mask-expansion detection. The former identifies subtle structural degradation through foreground-ratio attenuation, while the latter mitigates invalid mask growth by analyzing multi-frame ratio evolution. Given the scarcity of high-quality, densely annotated underwater video datasets, we develop a comprehensively annotated underwater filter-net video segmentation dataset, UWFN. Experimental results demonstrate that our proposed approach achieves a 𝒥&ℱ score of 88.2 on the UWFN dataset, exceeding classic methods and demonstrating superior robustness in real-world underwater inspection scenarios.

Jiawei Wang, Hongwen Yu, Zini Wang et al. · 0 citations
Open access Sep 2026

Underwater Computer Vision for Ecologists: A Framework for Curating Image Training Datasets for Object Detection

As marine ecosystems experience accelerating change, there is an urgent need for efficient and scalable biodiversity monitoring tools. We present a 10-step framework for integrating computer vision (CV) tools into long-term underwater biodiversity monitoring, using a case study from coastal British Columbia. Over 9000 h of unbaited remote underwater video footage were collected from two kelp farms and reference sites between March 2022 and June 2023. The framework includes steps for creating an annotated training dataset using unsupervised and supervised CV tools, culminating in the training and validation of a YOLOv8 object detection model. This process produced over 241,000 annotations across 54 pseudo-taxonomic categories (representing both taxa and visually similar groups of fauna), with a focus on fish and gelatinous zooplankton groups. The final model achieved an overall F1 (2 × precision × recall/(precision + recall)) of 0.74 and a mean average precision at 0.5 intersection over union (mAP50) of 0.78. The model had the highest performance on fine-resolution taxa such as Phanerodon vacca (F1 = 0.88) and Aurelia labiata (F1 = 0.90), and lowest performance on pseudo-taxa with limited visual distinctiveness such as Actinopterygii (F1 = 0.60) and Cnidaria (F1 = 0.60). A re-training experiment using annotation thresholds between 25 training images to full dataset availability (~200–24,000 images per group) found that model performance was positively correlated with annotation effort, with F1 averaging 0.84 and mAP50 averaging 0.91 at the maximum training dataset size. Our results suggest that the model is most effective for abundant and visually distinctive taxa, while performance declines for groups with coarse taxonomic resolution. We recommend optimizing annotation effort by targeting genus- or species-level taxa, having at least one broad-level group to capture order-level abundances, and supplementing annotations of rare groups with common but morphologically similar groups, which may further improve model performance.

Talen Rimmer, Colin Bates, Declan McIntosh et al. · 0 citations
Review Open access Aug 2026

WIO-ReefFish: A High-Resolution Dataset for Taxon-Aware Coral Reef Fish Detection in the Western Indian Ocean

WIO-ReefFish, a reef fish detection dataset derived from diver-operated line-intercept transects and designed for ecological monitoring under natural survey conditions, is presented and established as a realistic benchmark for automated reef fish detection and a foundation for more robust computer-vision tools in coral reef biodiversity monitoring.

J. Gerard, Luca Branger, F. Huyghe et al. · 0 citations
Open access Sep 2026

An underwater precision fish counting framework using transformer with feature offset aggregation and occlusion-aware attention in aquaculture

Accurate fish counting seeks to estimate the total number of fish within an image and is widely applied in fields such as sustainable aquaculture management, aquatic ecosystem monitoring, and population management. However, challenges arise due to variations in fish pose, occlusion resulting from aggregation, and background interference from elements such as aquatic plants, rocks, and suspended particles. To address these issues, this paper proposes a novel deep learning framework designed to enhance the accuracy and robustness of fish counting, termed OGLA-Net. First, Feature Offset Aggregation Strategy (FOAS) is employed to learn deformable representations to adapt to pose variations and mitigate inconsistencies in fish orientation, better coping with the variability of fish bodies during counting. Second, the Deformable-Guided Positional (DefoGP) module constructs an explicit guidance map in both spatial and frequency domains to focus on fish edges and group distributions, thereby enhancing the model’s capability to handle occlusions to solve the problem of dense fish counting. Finally, the Gated Soft Latent Attention(GSLA) mechanism suppresses redundant background textures through soft masking and adaptive gating of attention heads, effectively improving the anti-interference capability when monitoring complex underwater environments. The proposed method is systematically evaluated across three datasets: UGCD (uniform density distribution), CCD (complex background), and DGCD (high-density distribution). The method achieves MAE and RMSE values of 3.499 and 4.742 on the UGCD. It also maintained superior performance on the CCD and DGCD datasets. These results demonstrate that the proposed counting framework delivers high accuracy across diverse conditions, providing a reliable solution for automatic counting.

Tong-Tong Gu, Zheng-Meng Wu, Da-She Li et al. · 0 citations
Open access 2026

A Monocular Depth Estimation Framework for Improved Underwater Fish Biomass Assessment

Accurate and automated fish weight estimation is a critical component of modern aquaculture, enabling optimized feeding strategies, improved fish welfare, and effective production planning. A common approach is the use of computer vision techniques to estimate fish mass while the fish are swimming, thereby eliminating the need for manual handling and reducing stress on the fish. This paper presents a computer vision pipeline that integrates monocular depth estimation using MiDaS with YOLOv11 object detection, trained on a real-world underwater dataset of Nile tilapia covering multiple age groups. The new method requires only a monocular camera, eliminates the need for manual fish orientation, and enables fully automatic weight prediction through a CatBoost regressor. Experiments conducted across different fish age groups and tank conditions demonstrate consistent and robust performance. In this context, YOLOv11 achieves an F1-score of 0.93 at an IoU threshold of 0.5. Additionally, the weight estimation model attains a mean absolute error (MAE) below 4 g and an $R^{2}$ value exceeding 0.95. These results suggest that the proposed pipeline approach delivers accurate and reliable weight estimation without the need for using stereo-cameras, making it very well-suited for practical deployment in real-world aquaculture environments.

Said Al-Abri, S. Keshvari, Rami Al-Hmouz et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.