Skip to content
Preprint

TorchFDTD: Differentiable FDTD on a single GPU

Sep 2026 · 0 citations · 44 references
Physics

Abstract

Gradient-based photonic design needs full-wave derivatives with respect to millions of parameters, but on a single GPU workstation the time-domain adjoint is limited by device memory and by the incompatibility of fused update kernels with automatic differentiation. Here, we present TorchFDTD, an open-source finite-difference time-domain (FDTD) package that addresses both limits. Its Yee, absorber and dispersion updates execute as fused CUDA kernels captured in a CUDA graph, and every kernel is paired with a transpose kernel derived from its update, so material and geometry derivatives are obtained by a discrete adjoint that PyTorch chains with differentiable objectives built from the supported observations. For problems that exceed the device, a streamed mode advances the domain one causal slab at a time and keeps the global state and the checkpoints in host memory, which lowers the device allocation while preserving the resident discretization. We validate the package against analytic solutions, against Meep, FDTDX and a rigorous coupled-wave solver, and against automatic differentiation and finite differences. Against FDTDX on the same GPU its forward solves are 6.4 to 7.3 times faster and gradients 2.0 to 52 times faster, and in double precision on an A100 its solves are 49 to 59 times faster than Meep on a workstation CPU. On an RTX 3060, host streaming of $256^3$ and $320^3$ adjoints costs 3.1 and 2.7 times the resident time and lowers the peak device allocation by 56% and 65%. A 54-million-cell pillar-array lens coupled to an angular-spectrum objective yields an adjoint derivative within 0.78% of a central difference. The time-domain adjoint of a device with billions of cells thus becomes available on a single workstation GPU. Overlapping lateral tiles with angular-spectrum propagation evaluate a 1 mm $\times$ 1 mm metalens of $1.2\times10^{7}$ posts on one 48 GB GPU in 6.4 h.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.