Skip to content
Preprint

WildFin: An In-the-Wild Dataset for Fish Behavioral Recognition

Aug 2026 · 0 citations · 68 references
Computer Science

TL;DR

This work introduces WildFin, a novel benchmark for fish behavior recognition collected and annotated by ecologists, and benchmark modern vision foundation models and quantify tradeoffs between static and spatiotemporal architectures, revealing the substantial gap between current model capabilities and the demands of real-world underwater behavioral analysis.

Abstract

Recent advances in field technology have led to a massive influx of in-the-wild video data for ecological science. The primary bottleneck in leveraging this data is the high cost of expert annotation. While computer vision offers a potential solution, current models frequently fail when deployed in complex marine environments. To characterize these failures, we introduce WildFin, a novel benchmark for fish behavior recognition collected and annotated by ecologists. WildFin spans two critical real-world paradigms: stationary cameras monitoring groups of fish and dynamic divers following individual subjects. The dataset represents a massive curation effort, involving 1,350 hours of fieldwork and 600 hours of expert annotation to produce 9 hours of behavioral data with over 2 million frame-by-frame labels. We benchmark modern vision foundation models and quantify tradeoffs between static and spatiotemporal architectures, revealing the substantial gap that remains between current model capabilities and the demands of real-world underwater behavioral analysis. Project website: https://team-wildfin.github.io/.

View source

Similar papers

Review Open access Aug 2026

Bridging the edge–cloud gap: adaptive AI for robust image and audio wildlife monitoring

Artificial intelligence (AI) is transforming wildlife monitoring through automated analysis of images and acoustic recordings for tasks such as detection, filtering irrelevant events, and species identification. Many current approaches rely on large models trained on extensive datasets and deployed in the cloud, including systems such as MegaDetector or SpeciesNet for camera-trap imagery and BirdNET for avian acoustics. Although accurate under well-represented conditions, their performance often degrades when applied to new locations, species communities, or recording environments, highlighting persistent challenges in model generalization. Consequently, researchers increasingly rely on species- or site-specific models and adaptive strategies such as calibration, domain adaptation, and continual learning. At the same time, there is growing interest in moving computation closer to the sensor. Edge deployments using lightweight models on embedded platforms such as Raspberry Pi, Nvidia Jetson Nano, or AudioMoth enable real-time inference in remote environments, but introduce constraints related to computation, memory, and energy consumption. These trade-offs motivate hybrid edge-cloud architectures in which edge devices perform local filtering while more complex models and analysis remain in the cloud. This mini-review synthesizes advances in AI-based wildlife monitoring across image and audio modalities, focusing on generalization, data imbalance, and deployment on resource-constrained devices. We review emerging solutions including adaptive calibration, continual learning, and self-supervised representation learning, and discuss how multimodal AI and hybrid edge–cloud systems may enable scalable, robust, and context-aware ecological monitoring.

Delia Velasco-Montero, J. Fernández-Berni, R. Sankaran et al. · 0 citations
Review Open access Aug 2026

WIO-ReefFish: A High-Resolution Dataset for Taxon-Aware Coral Reef Fish Detection in the Western Indian Ocean

WIO-ReefFish, a reef fish detection dataset derived from diver-operated line-intercept transects and designed for ecological monitoring under natural survey conditions, is presented and established as a realistic benchmark for automated reef fish detection and a foundation for more robust computer-vision tools in coral reef biodiversity monitoring.

J. Gerard, Luca Branger, F. Huyghe et al. · 0 citations
Dataset Open access Aug 2026

A Dataset for Fish Segmentation and Tracking in Underwater Videos

An open-source, browser-based annotation tool integrating the Segment Anything Model (SAM2) and CUTIE for efficient semi-automatic segmentation and tracking and facilitates high-quality annotations without specialized hardware, improving accessibility and reproducibility within the marine imaging community.

Josep S. Sánchez, J. Lisani, I. A. Catalán et al. · 0 citations
Open access Aug 2026

DEEPFIN: A deep learning tool for fish image classification from unlabeled data

DEPFIN is an annotation-free image analysis pipeline that combines foreground segmentation, frozen pretrained convolutional encoders, non-linear dimensionality reduction, and density-based clustering for extracting biological structure from the growing volumes of unlabeled fish imagery.

Alexandru Mihai, Billy Moore, Marleen Klann et al. · 0 citations
Open access Aug 2026

Open-source tag-free monitoring of individual birds using automated weighing and deep-learning recognition

Effective animal monitoring is essential for assessing health, behavior, and environmental interactions, particularly in research and welfare contexts. This study presents a low-cost, open-source system designed for non-invasive monitoring of budgerigars (Melopsittacus undulatus), a small parrot species frequently used in animal behavior research. The system integrates a perch-based scale for voluntary weight measurement, a temperature sensor, and a camera for image capture, all controlled by a Raspberry Pi. By leveraging fine-tuned neural networks, the system achieves automated individual recognition with high accuracy, eliminating the need for invasive tagging methods. The modular design ensures accessibility, scalability, and minimal disturbance to the animals, while the accompanying software streamlines data collection, processing including labeling, and visualization. This approach provides a comprehensive solution for continuous monitoring, offering valuable insights for research and husbandry while prioritizing animal welfare.

Jinook Oh, Marisa Hoeschele · 0 citations
Open access Sep 2026

Underwater Computer Vision for Ecologists: A Framework for Curating Image Training Datasets for Object Detection

As marine ecosystems experience accelerating change, there is an urgent need for efficient and scalable biodiversity monitoring tools. We present a 10-step framework for integrating computer vision (CV) tools into long-term underwater biodiversity monitoring, using a case study from coastal British Columbia. Over 9000 h of unbaited remote underwater video footage were collected from two kelp farms and reference sites between March 2022 and June 2023. The framework includes steps for creating an annotated training dataset using unsupervised and supervised CV tools, culminating in the training and validation of a YOLOv8 object detection model. This process produced over 241,000 annotations across 54 pseudo-taxonomic categories (representing both taxa and visually similar groups of fauna), with a focus on fish and gelatinous zooplankton groups. The final model achieved an overall F1 (2 × precision × recall/(precision + recall)) of 0.74 and a mean average precision at 0.5 intersection over union (mAP50) of 0.78. The model had the highest performance on fine-resolution taxa such as Phanerodon vacca (F1 = 0.88) and Aurelia labiata (F1 = 0.90), and lowest performance on pseudo-taxa with limited visual distinctiveness such as Actinopterygii (F1 = 0.60) and Cnidaria (F1 = 0.60). A re-training experiment using annotation thresholds between 25 training images to full dataset availability (~200–24,000 images per group) found that model performance was positively correlated with annotation effort, with F1 averaging 0.84 and mAP50 averaging 0.91 at the maximum training dataset size. Our results suggest that the model is most effective for abundant and visually distinctive taxa, while performance declines for groups with coarse taxonomic resolution. We recommend optimizing annotation effort by targeting genus- or species-level taxa, having at least one broad-level group to capture order-level abundances, and supplementing annotations of rare groups with common but morphologically similar groups, which may further improve model performance.

Talen Rimmer, Colin Bates, Declan McIntosh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.