Micro-kernels for portable and efficient matrix multiplication in deep learning

Guillermo Alaejos,Enrique S Quintana-Ortí,Adrián Castelló,Francisco D Igual,Pedro Alonso-Jordá,Héctor Martínez

doi:10.1007/s11227-022-05003-3

Abstract

We provide a practical demonstration that it is possible to systematically generate a variety of high-performance micro-kernels for the general matrix multiplication (gemm) via generic templates which can be easily customized to different processor architectures and micro-kernel dimensions. These generic templates employ vector intrinsics to exploit the SIMD (single instruction, multiple data) units in current general-purpose processors and, for the particular type of gemm problems encountered in deep learning, deliver a floating-point throughput rate on par with or even higher than that obtained with conventional, carefully tuned implementations of gemm in current linear algebra libraries (e.g., BLIS, AMD AOCL, ARMPL). Our work exposes the structure of the template-based micro-kernels for ARM Neon (128-bit SIMD), ARM SVE (variable-length SIMD) and Intel AVX512 (512-bit SIMD), showing considerable performance for an NVIDIA Carmel processor (ARM Neon), a Fujitsu A64FX processor (ARM SVE) and on an AMD EPYC 7282 processor (256-bit SIMD).

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: The Journal of Supercomputing	Publication Date: Dec 14, 2022
Citations: 9	License type: open-access

R Discovery Prime

R Discovery Prime

Micro-kernels for portable and efficient matrix multiplication in deep learning

Abstract

Talk to us

Similar Papers

More From: The Journal of Supercomputing

Lead the way for us

Similar Papers

Enabling an OS kernel for large data with a SIMD Unit
Shogo Saito ... Shuichi Oikawa
-
Shogo Saito, et. al.Shogo Saito ... Shuichi Oikawa
01 Oct 2012
01 Oct 2012

Utilization of a SIMD Unit in the OS Kernel
Shogo Saito ... Shuichi Oikawa
-
Shogo Saito, et. al.Shogo Saito ... Shuichi Oikawa
01 Nov 2011
01 Nov 2011

Returning Control to the Programmer
Jonathan Parri ... Voicu Groza
Queue | VOL. 9
Jonathan Parri, et. al.Jonathan Parri ... Voicu Groza
01 Feb 2011
Queue | VOL. 9

Automatic Vectorization by Runtime Binary Translation
Takashi Nakamura ... Satoshi Miki
-
Takashi Nakamura, et. al.Takashi Nakamura ... Satoshi Miki
01 Nov 2011
01 Nov 2011

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Micro-kernels for portable and efficient matrix multiplication in deep learning

Abstract

Talk to us

Similar Papers

More From: The Journal of Supercomputing