Skip to content
Open access

OceanBench: A Benchmark for Data-Driven Global Ocean Forecasting systems

2025 · Advances in Neural Information Processing Systems 38 · 7 citations · ⚡ 2 influential · 32 references

TL;DR

OceanBench is a benchmark designed to evaluate and accelerate global short-range data-driven ocean forecasting, constructed from a curated dataset comprising first-guess trajectories, nowcasts, and atmospheric forcings from operational physical ocean models, typically unavailable in public datasets due to assimilation cycles.

Abstract

Data-driven approaches, particularly those based on deep learning, are rapidly advancing Earth system modeling. However, their application to ocean forecasting remains limited despite the ocean’s pivotal role in climate regulation and marine ecosystems. To address this gap, we present OceanBench, a benchmark designed to evaluate and accelerate global short-range (1-10 days) data-driven ocean forecasting. OceanBench is constructed from a curated dataset comprising first-guess trajectories, nowcasts, and atmospheric forcings from operational physical ocean models, typically unavailable in public datasets due to assimilation cycles. Matched observational data are also included, enabling realistic evaluation in an operational-like forecasting framework. The benchmark defines three complementary evaluation tracks: (i) Model-to-Reanalysis, where models are compared against the reanalysis dataset commonly used for training; (ii) Model-to-Analysis, assessing generalization to a higher-resolution physical analysis; and (iii) Model-to-Observations, Intercomparison and Validation (IV-TT) CLASS-4 evaluation against independent observational data. The first two tracks are further supported by process-oriented di-agnostics to assess the dynamical consistency and physical plausibility of forecasts. OceanBench includes key ocean variables: sea surface height, temperature, salinity, and currents, along with standardized metrics grounded in physical oceanography. Baseline comparisons with operational systems and state-of-the-art deep learning models are provided. All data, code, and evaluation protocols are openly available at https://github.com/mercator-ocean/oceanbench , establishing Ocean-Bench as a foundation for reproducible and rigorous research in data-driven ocean forecasting.

Read PDF

Similar papers

Open access Aug 2026

LangYa: a large AI model for global ocean forecasting.

Ocean forecasting is crucial for both scientific research and societal benefits. Large artificial intelligence (AI)-based models have recently boosted forecasting efficiency and accuracy. However, it remains challenging to develop a comprehensive AI-driven ocean forecasting system capable of integrating cross-spatiotemporal and atmospheric forcing. This study introduces LangYa, a cross-spatiotemporal and atmospheric forcing ocean forecasting system featuring: (1) a large-language-model-based (LLM-based) time embedding to explicitly represent forecast lead times, (2) an asynchronous cross-iterative random sampling strategy to represent the impacts of atmospheric forcing on ocean processes, (3) an ocean self-attention module to enhance network stability and accelerate training convergence, and (4) an adaptive loss function to capture ocean dynamics in the thermocline, at depths ranging from tens of meters to about 300 m. LangYa is trained on 27 years of global ocean data from the Global Ocean Reanalysis and Simulation, version 12 (GLORYS12). Using reanalysis and observational data, compared to existing open-source AI-based forecasting systems and numerical models, LangYa enables a single model to produce forecasts with lead times of 1 to 7 d (1/12°, daily) and achieves 7 d RMSEs below 0.0736 m/s, 0.0701 m/s, 0.4376 ℃, and 0.1302 psu for global currents, temperature, and salinity respectively. These quantitative results indicate that LangYa provides clear advantages in forecast accuracy, lead-time robustness, and stability for global OSV forecasting, demonstrating its potential for real-time operational deployment.

Nan Yang, Chong Wang, Zimeng Zhao et al. · 0 citations
Preprint Aug 2026

Deep Learning-Based Statistical Downscaling of Sea Surface Temperature Using a Residual Corrective Neural Network

This study proposes a novel deep learning framework that uses a U-Net to generate an initial high-resolution SST estimate, which is subsequently refined using a residual corrective approach, and progressively refines initial U-Net predictions by incorporating dynamically scaled residuals at each step, enabling accurate capture of broad patterns and fine-grained features such as eddies and fronts.

Onkar Jadhav, Tim French, I. Janeković et al. · 1 citation
Preprint Aug 2026

Simple data fusion from several ocean and atmosphere hindcast models improves surface drifter trajectory prediction

Simulating the trajectory of surface drifters in the ocean matters for search and rescue, pollution tracking, oil and chemical spill response, and marine risk analysis. Accurate prediction remains difficult, as widely acknowledged in the literature, and also illustrated by the ``Forecasting Floats in Turbulence''challenge issued by the US Defense Advanced Research Projects Agency (DARPA) in 2021, and which ultimately led to this paper. The main source of error usually comes from uncertain ocean currents, while errors in wind forcing and object drift properties are often smaller [Dagestad and R\"ohrs, 2019]. Here, we use an open one-year dataset of Sofar Spotter trajectories together with several ocean and atmospheric hindcast products to test data-driven drift models at scale. We compare three approaches: i) a standard (baseline) drifter trajectory simulation based on one ocean model and one atmospheric model, ii) linear regression (LR) models that fuse all available predictors, and iii) neural networks (NN) using similar inputs. A simple LR model that combines all predictors performs equally well as the NN. Because LR is simpler, cheaper, and more robust, we retain it as the preferred approach. In 2-day trajectory prediction, this improves the Liu-Weisberg skill score by around 40\% relative to the baseline. These findings apply to hindcast mode; applying this methodology for forecast mode remains for future work.

J. Rabault, K. Dagestad, G. Hope · 0 citations
Preprint Aug 2026

DLESyM-Ocean: A Deep Learning Probabilistic Global Model for Simulating Present-Day Upper Ocean and Sea Ice

While AI has shown remarkable promise in atmospheric and meteorological forecasting, accurately simulating other components of the Earth system with AI remains an active frontier. We present DLESyM-Ocean, a Deep Learning Earth System Model that simulates global present-day sea ice and upper ocean conditions. Unlike conventional probabilistic models optimized via diffusion objectives or losses such as continuous-ranked probability score, DLESyM-Ocean is trained using a patch energy score loss. When driven by atmospheric forcing, DLESyM-Ocean produces a well-calibrated, spatially coherent, and skillful ensemble of sea ice and upper ocean conditions with minimal bias relative to reanalysis products. DLESyM-Ocean is stable when autoregressively run for multi-year simulations and produces a climatology and variability with minimal bias compared with reanalysis. We evaluate case studies including a recent sea ice extreme, a severe marine heatwave, the 2023 El Ni\~no transition, and the 2023 spike in global mean temperature. In all of these case studies, DLESyM-Ocean produces realistic surface and subsurface trajectories and ample ensemble diversity in response to common atmospheric forcing, suggestive of learned autoregressive ocean dynamics. When coupled with other Earth system components, such as the atmosphere, the computational efficiency of DLESyM-Ocean makes it a promising tool for subseasonal to seasonal forecasting.

Zachary I. Espinosa, Nathaniel Cresswell-Clay, William Yik et al. · 0 citations
Preprint Jul 2026

Global reanalysis from observations alone with machine learning

Earth system reanalysis datasets are foundational for weather and climate research and provide the gridded training data used by most machine learning weather prediction systems. Here we show results from a prototype system that suggest that machine learning models trained only on Earth system observations can potentially be used to generate multi-decade global reanalyses without using physics-based numerical models. The resulting gridded fields capture large-scale atmospheric structure and variability across multiple timescales, while exhibiting signs of physical coherence in several key dynamical diagnostics. Evaluations of the prototype against held-out independent atmospheric observations indicate that the root mean square vector error of upper-level winds is close to that of ERA5 when compared at a consistent resolution, and that the standard deviation of the error at the surface is between that of 4th- and 5th-generation ECMWF reanalyses (ERA-Interim and ERA5). Furthermore, while traditional reanalysis production is computationally expensive, typically taking several years to produce, the reanalysis presented here was generated during the course of a single working day. These results suggest that observation-trained machine learning models offer a promising new approach for reanalysis production from observations alone.

Peter Lean, E. Pinnington, P. Laloyaux et al. · 0 citations
Preprint Jul 2026

Incomplete Observations Boost Evolutionary Performance in Ocean Modeling

Data-driven methods have revolutionized ocean modeling, yet current approaches rely heavily on complete reanalysis datasets, imposing computational constraints and limiting model performance to that of the training data. Here, we present a generative state-space model and an optimization framework that enable learning directly from sparse and noisy observations. The model is essentially a hidden Markov model with a continuous state space, where oceanic physical quantities are treated as hidden states and measurements as observations, enabling a unified representation of ocean fields and observational data. Both the initial-state and state-transition modules are implemented as neural networks to capture the complexity and temporal evolution of ocean states, while the emission module is formulated as a masked Gaussian distribution. To train the model from sparse observations, we derive an optimization framework based on the expectation-maximization (EM) algorithm. The framework alternately reconstructs high-fidelity ocean fields via Langevin dynamics and optimizes deep neural networks to capture temporal evolution. Theoretical analysis shows that the framework maximizes the likelihood of observations under the generative model. For efficiency, we assume that ocean-state evolution follows a stationary, ergodic, and Markovian stochastic process and adopt only length-two state sequences during optimization. Experiments on CMIP6 simulation data and FY-3D satellite data demonstrate high-fidelity reconstruction and accurate prediction, showing that sparse observations can directly improve the model's representation of ocean-state dynamics. This work offers a scalable pathway for next-generation Earth system models to learn directly from sparse, incomplete real-world observations.

Yangyang Kong, Yutong Jiang, Yanhai Gan et al. · 0 citations