Mixed precision LU factorization on GPU tensor cores: reducing data movement and memory footprint

Florent Lopez,Theo Mary

doi:10.1177/10943420221136848

Abstract

Modern GPUs equipped with mixed precision tensor core units present great potential to accelerate dense linear algebra operations such as LU factorization. However, state-of-the-art mixed half/single precision LU factorization algorithms all require the matrix to be stored in single precision, leading to expensive data movement and storage costs. This is explained by the fact that simply switching the storage precision from single to half leads to significant loss of accuracy, forfeiting all accuracy benefits from using tensor core technology. In this article, we propose a new factorization algorithm that is able to store the matrix in half precision without incurring any significant loss of accuracy. Our approach is based on a left-looking scheme employing single precision buffers of controlled size and a mixed precision doubly partitioned algorithm exploiting tensor cores in the panel factorizations. Our numerical results show that compared with the state of the art, the proposed approach is of similar accuracy but with only half the data movement and memory footprint, and hence potentially much faster: it achieves up to 2× and 3.5× speedups on V100 and A100 GPUs, respectively.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Mixed precision LU factorization on GPU tensor cores: reducing data movement and memory footprint

Abstract

Talk to us

Similar Papers

More From: The International Journal of High Performance Computing Applications

Lead the way for us

Journal: The International Journal of High Performance Computing Applications	Publication Date: Jan 3, 2023
Citations: 1

Similar Papers

NVIDIA Tensor Core Programmability, Performance & Precision
Stefano Markidis ... Ivy Bo Peng
-
Stefano Markidis, et. al.Stefano Markidis ... Ivy Bo Peng
01 May 2018
01 May 2018

Efficient Mixed-Precision Matrix Factorization of the Inverse Overlap Matrix in Electronic Structure Calculations with AI-Hardware and GPUs.
Adela Habib ... Anders M N Niklasson
Journal of chemical theory and computation | VOL. -
Adela Habib, et. al.Adela Habib ... Anders M N Niklasson
13 Aug 2024
Journal of chemical theory and computation | VOL. -

Mixed Precision Block Fused Multiply-Add: Error Analysis and Application to GPU Tensor Cores
Pierre Blanchard ... Nicholas J Higham
SIAM Journal on Scientific Computing | VOL. 42
Pierre Blanchard, et. al.Pierre Blanchard ... Nicholas J Higham
01 Jan 2020
SIAM Journal on Scientific Computing | VOL. 42

Mixed precision applied on common mathematical procedures over GPU
Marcelo A Sudo ... ÁLvaro L Fazenda
-
Marcelo A Sudo, et. al.Marcelo A Sudo ... ÁLvaro L Fazenda
19 Oct 2022
19 Oct 2022

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Mixed precision LU factorization on GPU tensor cores: reducing data movement and memory footprint

Abstract

Talk to us

Similar Papers

More From: The International Journal of High Performance Computing Applications