Bridging physics and data: A review of machine learning approaches for turbulence closure modeling
Abstract
Turbulence closure modeling remains one of the central challenges in computational fluid dynamics. Over the past decade, machine learning (ML) has emerged as a promising paradigm to augment and, in some cases, replace classical turbulence models by leveraging high-fidelity simulation data and advanced neural network architectures. We argue that a single obstacle recurs across the field—a mismatch between the pointwise, offline objective on which data-driven closures are trained and the solver-coupled, multi-scale, statistically stationary way they must perform—and we trace it to intrinsic physical properties of turbulence (extreme multi-scale chaos, non-locality, and the absence of scale separation). Through this lens, we survey three established pillars: (i) ML-enhanced Reynolds-averaged Navier–Stokes (RANS) closures, (ii) ML-based subgrid-scale models for large-eddy simulation, and (iii) physics-informed approaches, notably physics-informed neural networks (PINNs). Their characteristic failure modes—RANS ill-conditioning, the a priori/a posteriori gap, and PINN spectral bias—are local signatures of this common root. Emerging paradigms—neural operators, graph neural networks, transformers, diffusion models, and foundation models—are assessed by whether they change the objective or merely the architecture, and we propose concrete protocols for testing whether learned fields respect the Navier–Stokes dynamics statistically rather than pointwise. We critically examine generalizability across flow configurations, interpretability, numerical stability under solver coupling, and the scarcity of high-fidelity data, and conclude that aligning data-driven methods with physical constraints—and the training objective with deployment—is the path toward reliable, generalizable, and computationally efficient turbulence models for engineering applications.