Machine learning‐based auto‐tuning for enhanced performance portability of OpenCL applications

Thomas L Falch,Anne C Elster

doi:10.1002/cpe.4029

Abstract

SummaryHeterogeneous computing, combining devices with different architectures such as CPUs and GPUs, is rising in popularity and promises increased performance combined with reduced energy consumption. OpenCL has been proposed as a standard for programming such systems and offers functional portability. However, it suffers from poor performance portability, because applications must be retuned for every new device. In this paper, we use machine learning‐based auto‐tuning to address this problem. Benchmarks are run on a random subset of the tuning parameter spaces, and the results are used to build a machine learning‐based performance model. The model can then be used to find interesting subspaces for further search. We evaluate our method using five image processing benchmarks, with tuning parameter space sizes up to 2.3 M, using different input sizes, on several devices, including an Intel i7 4771 (Haswell) CPU, an Nvidia Tesla K40 GPU, and an AMD Radeon HD 7970 GPU. We compare different machine learning algorithms for the performance model. Our model achieves a mean relative error as low as 3.8% and is able to find solutions on average only 0.29% slower than the best configuration in some cases, evaluating less than 1.1% of the search space. The source code of our framework is available at https://github.com/acelster/ML‐autotuning.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Machine learning‐based auto‐tuning for enhanced performance portability of OpenCL applications

Abstract

Talk to us

Similar Papers

More From: Concurrency and Computation: Practice and Experience

Lead the way for us

Journal: Concurrency and Computation: Practice and Experience	Publication Date: Dec 22, 2016
Citations: 22

Similar Papers

ImageCL: An image processing language for performance portability on heterogeneous systems
Thomas L Falch ... Anne C Elster
-
Thomas L Falch, et. al.Thomas L Falch ... Anne C Elster
01 Jul 2016
01 Jul 2016

ImageCL: Language and source‐to‐source compiler for performance portability, load balancing, and scalability prediction on heterogeneous systems
Thomas L Falch ... Anne C Elster
Concurrency and Computation: Practice and Experience | VOL. 30
Thomas L Falch, et. al.Thomas L Falch ... Anne C Elster
20 Dec 2017
Concurrency and Computation: Practice and Experience | VOL. 30

Machine Learning Based Auto-Tuning for Enhanced OpenCL Performance Portability
Thomas L Falch ... Anne C Elster
-
Thomas L Falch, et. al.Thomas L Falch ... Anne C Elster
01 May 2015
01 May 2015

Analyzing and improving performance portability of OpenCL applications via auto-tuning
James Price ... Simon Mcintosh-Smith
-
James Price, et. al.James Price ... Simon Mcintosh-Smith
16 May 2017
16 May 2017

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Machine learning‐based auto‐tuning for enhanced performance portability of OpenCL applications

Abstract

Talk to us

Similar Papers

More From: Concurrency and Computation: Practice and Experience