Skip to content
Preprint

Diff-NekRS: A Scalable Differentiable Framework for Multi-Timestep Solver-in-the-Loop Training

Sep 2026 · 0 citations · 37 references
Physics

TL;DR

Diff-NekRS is introduced, a scalable differentiable framework that embeds neural corrections directly in the GPU-accelerated NekRS incompressible-flow solver, and establishes a verified and scalable path for multi-timestep solver-in-the-loop training that improves coarse-grid trajectory accuracy while retaining a speed advantage over the high-order reference.

Abstract

Hybrid physics-machine-learning solvers improve under-resolved simulations by embedding trainable corrections into the time integration. During autoregressive inference, repeated solver-model interactions can amplify small errors, motivating multi-timestep solver-in-the-loop training. However, production solvers rarely expose the derivatives needed to backpropagate through such rollouts. We introduce Diff-NekRS, a scalable differentiable framework that embeds neural corrections directly in the GPU-accelerated NekRS incompressible-flow solver. NekRS computes the authoritative forward trajectory, a manually implemented exact discrete adjoint differentiates the supported fully discrete timestep, and LibTorch supplies neural vector-Jacobian products and parameter gradients. End-to-end Taylor and centered finite-difference tests verify the assembled gradient for two-dimensional cylinder flow (2Dcyl) and the three-dimensional Taylor-Green vortex (3DTGV) across five horizons and 12-1,020 MPI ranks. At 1,020 ranks, optimizer-enabled post-setup training updates retain 54.5%-78.0% and 80.7%-81.9% weak-scaling efficiency for 2Dcyl and 3DTGV, respectively. In 200-step autoregressive inference, the M = 50 model reduces the three-seed median terminal relative L2 velocity error by 59.2% for 2Dcyl and 12.1% for 3DTGV relative to the uncorrected coarse-grid P = 2 baseline, and retains wall-clock speedups of 5.38x and 2.49x, respectively, relative to the corresponding P = 7 configurations for equal simulated-time intervals. These results establish a verified and scalable path for multi-timestep solver-in-the-loop training that improves coarse-grid trajectory accuracy while retaining a speed advantage over the high-order reference

View source

Similar papers

Preprint Aug 2026

RECAST: A Machine-Learning Framework for Correction and Super-Resolution of Coarse-Grid PDE Solvers

Results demonstrate that the learned correction and reconstruction capabilities of RECAST can enable substantially coarser PDE evolution without the corresponding loss of solution fidelity, providing a proof-of-concept route toward machine-learning acceleration of higher-dimensional numerical simulations across science...

M. Reza, F. Faraji · 0 citations
#artificial intelligence Preprint Sep 2026

Transolver-$\sigma$: Joint Spectral-Physical Subspace Modeling for Neural PDE Solving

Across five well-established PDE benchmarks spanning steady-state prediction and time-dependent dynamics, Transolver$ achieves state-of-the-art with a benchmark-averaged relative error reduction of 33.4% over the strongest baseline for each metric, while consistently improving autoregressive rollout over single-operato...

Haonan Shangguan, Hang Zhou, Haixu Wu et al. · 0 citations
Preprint Aug 2026

One-Step Evolution for Long-Time Extrapolation: An Error-Bound-Informed and Prior-Guided Neural Residual Framework for Autonomous PDEs

A numerical-prior-guided, physics-constrained method trained without ground-truth trajectory supervision that reduces long-time extrapolation error relative to the numerical prior and outperforms the best competing baseline in each case, thereby improving long-time simulation accuracy across different PDEs without grou...

Ma-Qun Zhang, Feng Gao, Wan-Kun Chen et al. · 0 citations
Conference Aug 2026

A Comprehensive Evaluation of Timestep Discretization Strategies in Text-to-Image Diffusion Models

Text-to-image latent diffusion models produce unprecedented visual fidelity but remain severely bottlenecked by the computational latency of iterative sampling. While optimizing the discretization of the continuous-time variable offers a powerful, training-free acceleration pathway, the comparative tradeoffs of foundat...

Tan-Yan Bao, Huy-Tan Thai · 0 citations
Preprint Aug 2026

From Numerical Simulators of PDEs to Neural Emulators and Back

Simulation is central to modern engineering and science, but the cost of numerical solvers for partial differential equations (PDEs) remains a bottleneck whenever fast or many-query evaluations are required. Neural emulators trained on solver-generated data promise significant speedups, yet they are usually framed as o...

Felix Koehler · 0 citations
#machine learning Preprint Aug 2026

RAPTOR: RAndom-projection Physics-informed Transient sOlveR

The complexity of time-domain simulation of modern power systems has increased significantly because converter-based resources introduce control dynamics that must be simulated alongside slower system-level and fast electromagnetic dynamics. The resulting wide range of timescales may force classical time-domain solvers...

Petros Ellinas, Benjamin Vilmann, Spyros Chatzivasileiadis et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.