The development of Mellanox/NVIDIA GPUDirect over InfiniBand—a new model for GPU to GPU communications

Gilad Shainer,Christian R Trott,Ali Ayoub,Paul S Crozier,Pak Lui,Michael Kagan,Greg Scantlen,Tong Liu

doi:10.1007/s00450-011-0157-1

Abstract

The usage and adoption of General Purpose GPUs (GPGPU) in HPC systems is increasing due to the unparalleled performance advantage of the GPUs and the ability to fulfill the ever-increasing demands for floating points operations. While the GPU can offload many of the application parallel computations, the system architecture of a GPU-CPU-InfiniBand server does require the CPU to initiate and manage memory transfers between remote GPUs via the high speed InfiniBand network. In this paper we introduce for the first time a new innovative technology--GPUDirect that enables Tesla GPUs to transfer data via InfiniBand without the involvement of the CPU or buffer copies, hence dramatically reducing the GPU communication time and increasing overall system performance and efficiency. We also explore for the first time the performance benefits of GPUDirect using Amber and LAMMPS applications.

Full Text