Modern Multi Research Articles

In fork-join parallelism , a sequential program is split into a directed acyclic graph of tasks linked by directed dependency edges, and the tasks are executed, possibly in parallel, in an order consistent with their dependencies. A popular and effective way to extend fork-join parallelism is to allow threads to create {futures . A thread creates a future to hold the results of a computation, which may or may not be executed in parallel. That result is returned when some thread touches that future, blocking if necessary until the result is ready. Recent research has shown that while futures can, of course, enhance parallelism in a structured way, they can have a deleterious effect on cache locality. In the worst case, futures can incur Ω(P T∞ + t T∞) deviations, which implies Ω(C P T∞ + C t T∞) additional cache misses, where C is the number of cache lines, P is the number of processors, t is the number of touches, and T∞ is the computation span . Since cache locality has a large impact on software performance on modern multicores, this result is troubling. In this paper, however, we show that if futures are used in a simple, disciplined way, then the situation is much better: if each future is touched only once, either by the thread that created it, or by a later descendant of the thread that created it, then parallel executions with work stealing can incur at most O(C P T 2 ∞) additional cache misses, a substantial improvement. This structured use of futures is characteristic of many (but not all) parallel applications.

Read full abstract

Current processor trends of integrating more cores with wider SIMD units, along with a deeper and complex memory hierarchy, have made it increasingly more challenging to extract performance from applications. It is believed by some that traditional approaches to programming do not apply to these modern processors and hence radical new languages must be discovered. In this paper, we question this thinking and offer evidence in support of traditional programming methods and the performance-vs-programming effort effectiveness of common multi-core processors and upcoming many-core architectures in delivering significant speedup, and close-to-optimal performance for commonly used parallel computing workloads. We first quantify the extent of the " Ninja gap ", which is the performance gap between naively written C/C++ code that is parallelism unaware (often serial) and best-optimized code on modern multi-/many-core processors. Using a set of representative throughput computing benchmarks, we show that there is an average Ninja gap of 24X (up to 53X ) for a recent 6-core Intel® Core™ i7 X980 Westmere CPU, and that this gap if left unaddressed will inevitably increase. We show how a set of well-known algorithmic changes coupled with advancements in modern compiler technology can bring down the Ninja gap to an average of just 1.3X . These changes typically require low programming effort, as compared to the very high effort in producing Ninja code. We also discuss hardware support for programmability that can reduce the impact of these changes and even further increase programmer productivity. We show equally encouraging results for the upcoming Intel® Many Integrated Core architecture (Intel® MIC) which has more cores and wider SIMD. We thus demonstrate that we can contain the otherwise uncontrolled growth of the Ninja gap and offer a more stable and predictable performance growth over future architectures, offering strong evidence that radical language changes are not required.

Read full abstract

Modern Multi Research Articles

Articles published on Modern Multi

Well-structured futures and cache locality

Bathymetry and geographical regionalization of Brepollen (Hornsund, Spitsbergen) based on bathymetric profiles interpolations

Can traditional programming bridge the Ninja performance gap for parallel computing applications?

Computational geometry in the parallel external memory model

UML Profile and Extensions for Complex Approval Systems with Complementary Levels of Abstraction

Effective Implementation of DGEMM on Modern Multicore CPU

Determination of optimal ordering quantity and reduction of bullwhip effect in a multistage supply chain using genetic algorithm

A Cost-Effective Hardware Approach for Measuring Power Consumption of Modern Multi-Core Processors

A simulation suite for Lattice-Boltzmann based real-time CFD applications exploiting multi-level parallelism on modern multi- and many-core architectures

Optimization Strategy of Top-Down Join Enumeration on Modern Multi-Core CPUs

Research of Collectives Optimization on Modern Multicore Clusters

A fast‐adaptive composite grid algorithm for solving the free‐space Poisson problem on the cell broadband engine

RAID Architecture with Correction of Corrupted Data in Faulty Disk Blocks

MUDD

Motivations and Performance Conditions for Ethnic Entrepreneurship

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Modern Multi Research Articles

Articles published on Modern Multi

Well-structured futures and cache locality

Bathymetry and geographical regionalization of Brepollen (Hornsund, Spitsbergen) based on bathymetric profiles interpolations

Can traditional programming bridge the Ninja performance gap for parallel computing applications?

Computational geometry in the parallel external memory model

UML Profile and Extensions for Complex Approval Systems with Complementary Levels of Abstraction

Effective Implementation of DGEMM on Modern Multicore CPU

Determination of optimal ordering quantity and reduction of bullwhip effect in a multistage supply chain using genetic algorithm

A Cost-Effective Hardware Approach for Measuring Power Consumption of Modern Multi-Core Processors

A simulation suite for Lattice-Boltzmann based real-time CFD applications exploiting multi-level parallelism on modern multi- and many-core architectures

Optimization Strategy of Top-Down Join Enumeration on Modern Multi-Core CPUs

Research of Collectives Optimization on Modern Multicore Clusters

A fast‐adaptive composite grid algorithm for solving the free‐space Poisson problem on the cell broadband engine

RAID Architecture with Correction of Corrupted Data in Faulty Disk Blocks

MUDD

Motivations and Performance Conditions for Ethnic Entrepreneurship