Aug 2026· ACM Transactions on Architecture and Code Optimization (TACO)· 0 citations· 34 references
TL;DR
FNO-Speed, an integrated solution incorporating the multi-level parallel FNO-aware mapping and tiling GEMM optimization strategy and the custom-sized high-frequency signal filtering scheme, is proposed, fully demonstrating the effectiveness of the FNO-Speed optimization strategy in improving FNO performance.
Abstract
Deep learning for solving partial differential equations (PDEs) has become increasingly prominent. The Fourier Neural Operator (FNO) architecture has been proven to be an efficient and high-precision method that is widely used in scientific research. However, FNO incurs significant overhead by increasing the scale and dimensionality of practical problems. The insufficient utilization of hardware resources in its key operations reduces the computational efficiency of FNO solvers in high-resolution and time-sensitive problem scenarios, and cannot provide effective solution capabilities. To address the latency induced by low computational resource utilization and large-scale data access and computation, we propose FNO-Speed, an integrated solution incorporating the multi-level parallel FNO-aware mapping and tiling GEMM optimization strategy and the custom-sized high-frequency signal filtering scheme. FNO-Speed effectively leverages the data characteristics of FNO layers and the GPU hierarchical structure to adopt a data tiling and partitioning strategy, implementing matrix multiplication based on vector outer products and operator fusion to replace convolution. It also adopts a data reorganization scheme and computation restructuring to address fragmented memory access operations and serial einsum in frequency-domain. The FNO-Speed optimization strategy enhances the utilization of device memory bandwidth and computational efficiency and achieves significant acceleration in both 2D and 3D problem scenarios while maintaining nearly identical accuracy. The model achieves up to 1.4 × end-to-end training speedup, and the parallel efficiency achieves around 70% on 4 GPUs, fully demonstrating the effectiveness of the FNO-Speed optimization strategy in improving FNO performance.
Neural operators like the Fourier Neural Operator (FNO) have demonstrated remarkable success in solving partial differential equations (PDEs) but suffer from high training costs due to fast Fourier Transform (FFT) operations. Targeting the unique computational bottleneck of the FFT in FNO, FNO–MP achieves acceleration...
Yi-Yang Zhu· International Conference on...· 0 citations
Two general multi-stage neural operator learning frameworks applicable when the target operator can be represented by a PDE, leveraging the weak form of the PDE residual for training are introduced.
Zhiping Mao, Zhenye Wen, Yong Zhang et al.· 0 citations
FlashPDE provides a hardware-efficient execution layer that bridges differentiable PDE solvers and GPU-optimized numerical computation within the PyTorch ecosystem, while maintaining numerical agreement with PyTorch finite-difference references.
Field temporal prediction and source identification constitute canonical problems in dynamical systems. Conventional approaches to these problems depend on a thorough understanding of the governing partial differential equations (PDEs). Recently, deep learning, as represented by neural operators, has provided a data-dr...
Bai-Ming Zhang, Jin-Song Tang, Ying Xu et al.· 0 citations
Edge computing and artificial intelligence have made the efficient deployment of machine vision algorithms on low-power hardware a critical challenge for integrated circuit design. Given data-intensive image pixels and deep neural network tensors, traditional von Neumann architectures inevitably encounter severe memory...
Linenxu Zhang· MATEC Web of Conferences· 0 citations
Overall, query-dependent cross-attention is the most reliable mechanism, whereas branch self-attention is most useful for large, spatially complex functional inputs, whereas branch self-attention is most useful for large, spatially complex functional inputs.
Amar Alem Koric, Qi-Bang Liu, S. Koric· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.