An Optimization for MapReduce Frameworks in Multi-core Architectures

Tharso Ferreira,Antonio Espinosa,Juan Carlos Moure,Porfidio Hernández

doi:10.1016/j.procs.2013.05.446

Tharso Ferreira, Antonio Espinosa + Show 2 more

Open Access

https://doi.org/10.1016/j.procs.2013.05.446

Copy DOI

Journal: Procedia Computer Science	Publication Date: Jan 1, 2013
Citations: 6	License type: cc-by-nc-nd

Affiliation: Autonomous University of Barcelona

Abstract

MapReduce simplifies parallel programming, abstracting the programmer responsibilities as synchronization and task man- agement. The paradigm allows the programmer to write sequential code which is automatically parallelized. The MapReduce frameworks developed today are designed for situations where all keys generated by the Map phase must fit into main memory. However certain types of workload have a distribution of keys that provoke a growth of intermediate data structures, exceeding the amount of available main memory. Based on the behavior of MapReduce frameworks in multi-core architectures for these types of workload, we promote an extension of the original strategy of MapReduce for multi-core architectures. We present an extension in memory hierarchy, hard disk and main memory, which has as objective to reduce the use of main memory, as well as reducing the page faults, caused by the use of swap. The main goal of our extension is to ensure an acceptable performance of MapReduce, when intermediate data structures do not fit in main memory and it is necessary to make use of a secondary memory.

Full Text