A hybrid CPU/GPU approach for optimizing sorting throughput

Michael Gowanlock,Ben Karsin

doi:10.1016/j.parco.2019.01.004

Abstract

The GPU is an effective architecture for sorting due to its massive parallelism and high memory bandwidth. However, for input datasets that exceed global memory capacity, the communication overhead between host (CPU) and GPU may degrade the overall performance of heterogeneous approaches. Thus, to achieve performance gains over multi-core parallel CPU algorithms, heterogeneous sorting using the GPU needs to obviate communication overheads. We provide a detailed overview of current host-GPU data transfer mechanisms and advance several methods of mitigating the associated performance bottlenecks. Using these methods, we develop a heterogeneous CPU/GPU sorting algorithm that effectively exploits the architecture. Furthermore, we demonstrate that, while out-of-place GPU sorting achieves the best performance, an in-place sort has the potential to further reduce some host-side bottlenecks, which encourages several future research priorities. Our approaches mitigate several bottlenecks, as demonstrated on single- and dual-GPU platforms, achieving speedups up to 3.47× over the parallel reference implementation on the CPU. We discuss future research for heterogeneous sorting in the multi-GPU NVLink era.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

A hybrid CPU/GPU approach for optimizing sorting throughput

Abstract

Talk to us

Similar Papers

More From: Parallel Computing

Lead the way for us

Journal: Parallel Computing	Publication Date: Feb 1, 2019
Citations: 12

Similar Papers

A distributed in-memory key-value store system on heterogeneous CPU–GPU cluster
Kai Zhang ... Bei Hua
The VLDB Journal | VOL. 26
Kai Zhang, et. al.Kai Zhang ... Bei Hua
21 Aug 2017
The VLDB Journal | VOL. 26

Two-level main memory co-design: Multi-threaded algorithmic primitives, analysis, and simulation
Michael A Bender ... Arun Rodrigues
Journal of Parallel and Distributed Computing | VOL. 102
Michael A Bender, et. al.Michael A Bender ... Arun Rodrigues
03 Jan 2017
Journal of Parallel and Distributed Computing | VOL. 102

Feasibility of a best-worst scaling exercise to set priorities for autism research.
Scott A Davis ... Daniel E Jonas
Health expectations : an international journal of public participation in health care and health policy | VOL. 25
Scott A Davis, et. al.Scott A Davis ... Daniel E Jonas
08 Jun 2022
Health expectations : an international journal of public participation in health care and health policy | VOL. 25

Mega-KV
Kai Zhang ... Xiaodong Zhang
Proceedings of the VLDB Endowment | VOL. 8
Kai Zhang, et. al.Kai Zhang ... Xiaodong Zhang
01 Jul 2015
Proceedings of the VLDB Endowment | VOL. 8

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A hybrid CPU/GPU approach for optimizing sorting throughput

Abstract

Talk to us

Similar Papers

More From: Parallel Computing