Benchmarking the Memory Hierarchy of Modern GPUs

Xinxin Mei,Kaiyong Zhao,Xiaowen Chu,Chengjian Liu

doi:10.1007/978-3-662-44917-2_13

Abstract

Memory access efficiency is a key factor for fully exploiting the computational power of Graphics Processing Units (GPUs). However, many details of the GPU memory hierarchy are not released by the vendors. We propose a novel fine-grained benchmarking approach and apply it on two popular GPUs, namely Fermi and Kepler, to expose the previously unknown characteristics of their memory hierarchies. Specifically, we investigate the structures of different cache systems, such as data cache, texture cache, and the translation lookaside buffer (TLB). We also investigate the impact of bank conflict on shared memory access latency. Our benchmarking results offer a better understanding on the mysterious GPU memory hierarchy, which can help in the software optimization and the modelling of GPU architectures. Our source code and experimental results are publicly available.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Benchmarking the Memory Hierarchy of Modern GPUs

Abstract

Talk to us

Similar Papers

Lead the way for us

Publication Date: Jan 1, 2014
Citations: 57	License type: cc-by

Similar Papers

Dissecting GPU Memory Hierarchy Through Microbenchmarking
Xinxin Mei ... Xiaowen Chu
IEEE Transactions on Parallel and Distributed Systems | VOL. 28
Xinxin Mei, et. al.Xinxin Mei ... Xiaowen Chu
01 Jan 2017
IEEE Transactions on Parallel and Distributed Systems | VOL. 28

Massively Parallel Network Coding on GPUs
Xiaowen Chu ... Kaiyong Zhao
-
Xiaowen Chu, et. al.Xiaowen Chu ... Kaiyong Zhao
01 Dec 2008
01 Dec 2008

STD-TLB: A STT-RAM-based dynamically-configurable translation lookaside buffer for GPU architectures
Xiaoxiao Liu ... Alex K Jones
-
Xiaoxiao Liu, et. al.Xiaoxiao Liu ... Alex K Jones
01 Jan 2014
01 Jan 2014

A Hierarchically Blocked Jacobi SVD Algorithm for Single and Multiple Graphics Processing Units
Vedran Novaković
SIAM Journal on Scientific Computing | VOL. 37
Vedran NovakovićVedran Novaković
01 Jan 2015
SIAM Journal on Scientific Computing | VOL. 37

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Benchmarking the Memory Hierarchy of Modern GPUs

Abstract

Talk to us

Similar Papers