FlashPDE provides a hardware-efficient execution layer that bridges differentiable PDE solvers and GPU-optimized numerical computation within the PyTorch ecosystem, while maintaining numerical agreement with PyTorch finite-difference references.
Abstract
Physics-Informed Neural Networks (PINNs) solve PDEs by incorporating physical constraints into neural-network training, but large-scale problems are limited by automatic-differentiation memory overhead and inefficient execution of grid-based PDE operators. We present FlashPDE, a drop-in fused operator library for grid-based scientific machine learning. FlashPDE replaces fragmented PyTorch finite-difference execution with differentiable Triton kernels. Each operator integrates fused stencil evaluation, an analytic discrete-adjoint backward pass, and boundary-gradient correction within a unified PyTorch autograd Function interface. The library provides 14 differentiable PDE operators covering 17 configurations across 1D--3D elliptic, parabolic, and Navier--Stokes systems, while remaining independent of neural architectures and training strategies. Experiments on an NVIDIA A100 GPU show that FlashPDE reduces peak memory usage by up to 37.0x compared with coordinate-based automatic differentiation and reduces CUDA kernel launches by up to 3.5x compared with eager PyTorch finite-difference implementations. Across six representative PDE benchmarks, FlashPDE achieves up to 2.30x end-to-end time-to-solution speedup and up to 19.2x kernel-level acceleration while maintaining numerical agreement with PyTorch finite-difference references. FlashPDE provides a hardware-efficient execution layer that bridges differentiable PDE solvers and GPU-optimized numerical computation within the PyTorch ecosystem.
Simulation is central to modern engineering and science, but the cost of numerical solvers for partial differential equations (PDEs) remains a bottleneck whenever fast or many-query evaluations are required. Neural emulators trained on solver-generated data promise significant speedups, yet they are usually framed as o...
FNO-Speed, an integrated solution incorporating the multi-level parallel FNO-aware mapping and tiling GEMM optimization strategy and the custom-sized high-frequency signal filtering scheme, is proposed, fully demonstrating the effectiveness of the FNO-Speed optimization strategy in improving FNO performance.
Tao Song, Ying Ouyang, Xiangyu Meng et al.· ACM Transactions on Architec...· 0 citations
Physics-informed neural networks (PINNs) and neural operators are increasingly used to learn solutions and solution operators for partial differential equations, but the literature remains fragmented across architectures, training objectives, and benchmark protocols. This review organizes these developments for a neura...
Namhyeon Kim, Mukhiddin Toshpulatov, Suan Lee· Neural Networks· 0 citations
Hybrid physics-machine-learning solvers improve under-resolved simulations by embedding trainable corrections into the time integration. During autoregressive inference, repeated solver-model interactions can amplify small errors, motivating multi-timestep solver-in-the-loop training. However, production solvers rarely...
Junoh Jung, R. Balin, Bethany Lusch et al.· 0 citations
We systematically investigate finite-difference (FD) derivative computation in Physics-Informed Neural Networks (PINNs) as an alternative to automatic differentiation (AD). On three benchmark PDEs we show that, with a properly calibrated step size, FD matches AD in accuracy on every problem while running faster across...
Gradient-based photonic design needs full-wave derivatives with respect to millions of parameters, but on a single GPU workstation the time-domain adjoint is limited by device memory and by the incompatibility of fused update kernels with automatic differentiation. Here, we present TorchFDTD, an open-source finite-diff...
Hyoseok Park· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.