Optimizing Finite Volume Method Solvers on Nvidia GPUs

Jingheng Xu,Conghui He,Guangwen Yang,Wayne Luk,Chao Yang,Wen Shi,Haohuan Fu,Yong Jiang,Wei Xue,Lin Gan

doi:10.1109/tpds.2019.2926084

Abstract

As scientific applications are increasingly ported to GPUs to benefit from both the powerful computing capacity and high throughput, accelerating explicit solvers for GPU-based finite volume methods is gaining more and more attention. In this paper, based on the detailed analysis of the FVM algorithm, we present a set of novel optimization methods, including the explicit data cache mechanism, optimal global memory loading strategy, as well as the inner-thread rescheduling method, which derives a suitable mapping from the solver algorithm to the underlying GPU hardware architecture, so as to remarkably improve the solving performance of structured mesh based FVM. We demonstrate the impact of our tuning techniques on two widely-used atmospheric dynamic kernels (3-D Euler and 2-D SWE) on five kinds of mainstream GPU platforms, and make a detailed analysis of the different tuning methodologies so as to demonstrate how to select the proper tuning strategy to different applications on various GPU platforms. Specifically, 93.9x speedup is achieved for the 3D Euler solver on Nvidia V100 over one 12-core Intel E5-2697 (v2) CPU, which is a 77 percent improvement compared with the original speedup without adopting the tuning techniques presented in this work.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Optimizing Finite Volume Method Solvers on Nvidia GPUs

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Parallel and Distributed Systems

Lead the way for us

Journal: IEEE Transactions on Parallel and Distributed Systems	Publication Date: Dec 1, 2019
Citations: 40

Similar Papers

Generalized GPU Acceleration for Applications Employing Finite-Volume Methods
Jingheng Xu ... Wei Xue
-
Jingheng Xu, et. al.Jingheng Xu ... Wei Xue
01 May 2016
01 May 2016

High Throughput JPEG 2000 for Video Content Production and Delivery Over IP Networks
David Taubman ... Reji Mathew
Frontiers in Signal Processing | VOL. 2
David Taubman, et. al.David Taubman ... Reji Mathew
27 Apr 2022
High Throughput JPEG 2000 for Video Content Production and Delivery Over IP Networks
David Taubman ... Reji Mathew

Investigation of lumbar spine biomechanics using global convergence optimization and constant loading path methods.
Won Man Park ... Young Joon Kim
Mathematical Biosciences and Engineering | VOL. 17
Won Man Park, et. al.Won Man Park ... Young Joon Kim
01 Jan 2020
Mathematical Biosciences and Engineering | VOL. 17

Particle Swarm Optimization Applied to Spacecraft Reentry Trajectory
Afshin Rahimi ... Hekmat Alighanbari
Journal of Guidance, Control, and Dynamics | VOL. 36
Afshin Rahimi, et. al.Afshin Rahimi ... Hekmat Alighanbari
18 Dec 2012
Journal of Guidance, Control, and Dynamics | VOL. 36

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Optimizing Finite Volume Method Solvers on Nvidia GPUs

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Parallel and Distributed Systems