Compressed basis GMRES on high-performance graphics processing units

José I Aliaga,Hartwig Anzt,Thomas Grützmacher,Enrique S Quintana-Ortí,Andrés E Tomás

doi:10.1177/10943420221115140

Abstract

Krylov methods provide a fast and highly parallel numerical tool for the iterative solution of many large-scale sparse linear systems. To a large extent, the performance of practical realizations of these methods is constrained by the communication bandwidth in current computer architectures, motivating the investigation of sophisticated techniques to avoid, reduce, and/or hide the message-passing costs (in distributed platforms) and the memory accesses (in all architectures). This article leverages Ginkgo’s memory accessor in order to integrate a communication-reduction strategy into the (Krylov) GMRES solver that decouples the storage format (i.e., the data representation in memory) of the orthogonal basis from the arithmetic precision that is employed during the operations with that basis. Given that the execution time of the GMRES solver is largely determined by the memory accesses, the cost of the datatype transforms can be mostly hidden, resulting in the acceleration of the iterative step via a decrease in the volume of bits being retrieved from memory. Together with the special properties of the orthonormal basis (whose elements are all bounded by 1), this paves the road toward the aggressive customization of the storage format, which includes some floating-point as well as fixed-point formats with mild impact on the convergence of the iterative process. We develop a high-performance implementation of the “compressed basis GMRES” solver in the Ginkgo sparse linear algebra library using a large set of test problems from the SuiteSparse Matrix Collection. We demonstrate robustness and performance advantages on a modern NVIDIA V100 graphics processing unit (GPU) of up to 50% over the standard GMRES solver that stores all data in IEEE double-precision.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: The International Journal of High Performance Computing Applications	Publication Date: Aug 5, 2022
Citations: 2	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Compressed basis GMRES on high-performance graphics processing units

Abstract

Talk to us

Similar Papers

More From: The International Journal of High Performance Computing Applications

Lead the way for us

Similar Papers

Multiple k −opt evaluation multiple k −opt moves with GPU high performance local search to large-scale traveling salesman problems
Wen-Bao Qiao ... Jean-Charles Créput
Annals of Mathematics and Artificial Intelligence | VOL. 88
Wen-Bao Qiao, et. al.Wen-Bao Qiao ... Jean-Charles Créput
01 Apr 2020
Annals of Mathematics and Artificial Intelligence | VOL. 88

Field Programmable Gate Arrays for Enhancing the Speed and Energy Efficiency of Quantum Dynamics Simulations.
José M Rodrı́Guez-Borbón ... Bryan M Wong
Journal of chemical theory and computation | VOL. 16
José M Rodrı́Guez-Borbón, et. al.José M Rodrı́Guez-Borbón ... Bryan M Wong
27 Mar 2020
Journal of chemical theory and computation | VOL. 16

Trident: A Hybrid Correlation-Collision GPU Cache Timing Attack for AES Key Recovery
Jaeguk Ahn ... John Kim
-
Jaeguk Ahn, et. al.Jaeguk Ahn ... John Kim
01 Feb 2021
01 Feb 2021

Work-in-Progress: NVIDIA GPU Scheduling Details in Virtualized Environments
Nicola Capodieci ... Roberto Cavicchioli
-
Nicola Capodieci, et. al.Nicola Capodieci ... Roberto Cavicchioli
01 Sep 2018
01 Sep 2018

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Compressed basis GMRES on high-performance graphics processing units

Abstract

Talk to us

Similar Papers

More From: The International Journal of High Performance Computing Applications