Skip to content
Review Open access

A Review of GPU-Accelerated Finite-Difference Numerical Simulation: Four-Phase Synthesis, Design Rules, and a Metrics Card

2026 · IEEE Access · Vol 14, pp. 120733-120761 · 0 citations · 80 references
Computer Science

Abstract

This review synthesizes research on graphics processing unit (GPU)-accelerated finite-difference numerical simulation (FDNS) from 2003 to 2025 to clarify how GPU computing has reshaped simulation workflows rather than merely accelerating isolated numerical kernels and to address challenges related to comparability, verification, communication, and reporting in production-oriented FDNS. It combines a Web of Science bibliometric corpus with structured extraction from representative studies reporting performance, scale, precision, accuracy, and implementation characteristics, and organizes the evidence into a four-phase evolution model: graphics-era prototyping, Compute Unified Device Architecture and Open Computing Language expansion and early multi-GPU deployment, systematic optimization, and exascale-adjacent maturity. The synthesis shows that problem scales have grown from roughly one-million-cell demonstrations to multi-GPU and cluster-scale workloads approaching 10 billion unknowns, while precision practice has shifted from single or reduced precision toward double-precision production runs and selectively validated mixed precision; identifies three workload profiles—bandwidth-bound local stencil time stepping, stencil pipelines coupled to global operators, and adaptive or heterogeneous workflows; derives six design rules for memory locality, single-instruction multiple-thread regularity, compute–communication overlap, adaptive work distribution, precision-aware verification, and workflow-level energy and input/output reporting; proposes a minimal metrics card covering the hardware/software stack, baseline, timing scope, precision policy, accuracy evidence, throughput, energy, and input/output inclusion; and outlines a staged modernization blueprint for legacy FDNS solvers. The review concludes that GPU-accelerated FDNS is now an end-to-end workflow problem requiring optimized stencil kernels, explicit numerical verification, transparent reporting, communication-aware orchestration, and reproducible treatment of precision, energy, and storage costs.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.