Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

A novel power model for future heterogeneous 3D chip-multiprocessors in the dark silicon age

  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Dark silicon has recently emerged as a new problem in VLSI technology. Maximizing performance of chip-multiprocessors (CMPs) under power and thermal constraints is very challenging in the dark silicon era. Providing next-generation analytical models for future CMPs which consider the impact of power consumption of core and uncore components such as cache hierarchy and on-chip interconnect that consume significant portion of the on-chip power consumption is largely unexplored. In this article, we propose a detailed power model which is useful for future CMP power modeling. In the proposed architecture for future CMPs, we exploit emerging technologies such as non-volatile memories (NVMs) and 3D techniques to combat dark silicon. Results extracted from the simulations are compared with those obtained from the analytical model. Comparisons show that the proposed model accurately estimates the power consumption of CMPs running both multi-threaded and multi-programed workloads.

Similar Papers
  • Research Article
  • Cite Count Icon 24
  • 10.1016/j.micpro.2017.03.011
Optimization-based power and thermal management for dark silicon aware 3D chip multiprocessors using heterogeneous cache hierarchy
  • Apr 14, 2017
  • Microprocessors and Microsystems
  • Arghavan Asad + 3 more

Optimization-based power and thermal management for dark silicon aware 3D chip multiprocessors using heterogeneous cache hierarchy

  • Book Chapter
  • Cite Count Icon 1
  • 10.1016/bs.adcom.2018.04.001
Revisiting Processor Allocation and Application Mapping in Future CMPs in Dark Silicon Era
  • Jan 1, 2018
  • Mohaddeseh Hoveida + 5 more

Revisiting Processor Allocation and Application Mapping in Future CMPs in Dark Silicon Era

  • Conference Article
  • Cite Count Icon 6
  • 10.1109/dsd.2015.42
Exploiting Heterogeneity in Cache Hierarchy in Dark-Silicon 3D Chip Multi-processors
  • Aug 1, 2015
  • Arghavan Asad + 3 more

Technology scaling has enabled increasing number of cores on a chip in Chip-Multiprocessors (CMPs). As the number of cores increases, the overall system will need to provide more cache resources to feed all the cores. However, increasing the size of each cache level in the cache hierarchy of CMPs mitigates the large off-chip memory access latencies and bandwidth constraints. Moreover, cache hierarchy is known as one of the most power-hungry components in many-core CMPs because leakage power within the cache systems has become a significant contributor in the overall chip power budget in deep sub-micron as well as dark silicon era. Due to the many advantages of Non-Volatile Memory (NVM) technology such as high density, near zero leakage, and non-volatility, in this paper, we focus on exploiting such memories in the cache hierarchy. Specifically, we focus on 3D CMPs to decrease the leakage power consumption and mitigating the dark silicon phenomenon. Experimental results show that the proposed method on average improves the throughput by 45.5% and energy-delay product by 56% when compared to the conventional single cache technology.

  • Conference Article
  • Cite Count Icon 10
  • 10.1109/recosoc.2015.7238085
Energy efficient 3D Hybrid processor-memory architecture for the dark silicon age
  • Jan 1, 2015
  • Sobhan Niknam + 3 more

With increasing the number of cores on the Chip-Multiprocessors (CMPs) as a result of continuous technology scaling, more cache resources are needed to feed all the cores. Hence, in order to improve performance by reducing off-chip memory access, inevitably on-chip caches should be increased. In on-chip cache hierarchy, last level cache (LLCs) is the largest one consuming more energy compared with the other levels in many-core CMPs as leakage power within the LLC has become a significant contributor in the overall chip power budget in deep sub-micron as well as dark silicon era. In this paper, we focus on exploiting Non-Volatile Memory (NVM) which is a new type of memory with promising features in shared distributed LLCs to decrease the leakage power consumption and mitigating the dark silicon phenomenon. In our proposed strategy, we first calculate Average Memory Access Time (AMAT) of running applications on the CMP in each predetermined interval by collected systems memory traffic. Based on the monitored AMATs, we then adaptively reconfigure Hybrid distributed LLC by selecting the proper memory type (i.e., SRAM bank or STT-RAM bank) at runtime. Experiment results on the PARSEC benchmarks show that the proposed method provides up to 55.22% (on average 39.3%) energy reduction and 35.33% on average energy-delay product (EDP) improvement with only 6% performance degradation compared to the conventional methods where single cache technology is used.

  • Conference Article
  • Cite Count Icon 1
  • 10.1109/mcsoc.2015.42
Lighting the Dark-Silicon 3D Chip Multi-processors by Exploiting Heterogeneity in Cache Hierarchy
  • Sep 1, 2015
  • Ashkan Sadeghi + 3 more

This paper addresses a set of design paradigms by exploiting device and architectural heterogeneity to mitigate the dark silicon. We exploit Non-Volatile Memory (NVM) as potential replacements to conventional caches. Also, we study the problem of dynamic thread mapping in future Chip Multi-Processors (CMPs) via an efficient scheduler. Evaluations on a 3D architecture consisting of 8 core (a big and a small core on each tile) show that the proposed method provides up to 7% average performance improvement for multithreaded benchmarks, and 9% for multiprogrammed workloads. The results also show 62.5% and 67.7% on average energy-delay product (EDP) improvement for multithreaded and multiprogrammed workloads respectively, with 7.87% area overhead compared to the conventional methods.

  • Book Chapter
  • 10.1007/978-3-319-31596-6_8
Robust Application Scheduling with Adaptive Parallelism in Dark-Silicon Constrained Multicore Systems
  • Jan 1, 2017
  • Nishit Kapadia + 1 more

With deeper technology scaling accompanied by a worsening power-wall, an increasing proportion of chip area on a chip multiprocessor (CMP) is expected to be occupied by dark silicon. At the same time, design challenges due to process variations and soft-errors in integrated circuits are projected to become even more severe. It is well known that spatial variations in process parameters introduce significant unpredictability in the performance and power profiles of CMP cores. By mapping applications on to the best set of cores, process variations could potentially be used to our advantage in the dark-silicon era. In addition, the probability of occurrence of soft-errors during execution of any application has been found to be strongly related to the supply voltage and operating frequency values, thus necessitating reliability awareness within run-time voltage scaling schemes in contemporary CMPs. In this chapter, we present a novel framework that leverages the knowledge of variations on the chip to perform run-time application mapping and dynamic voltage scaling (DVS) to optimize system performance and energy, while satisfying dark-silicon power constraints of the chip as well as application-specific performance and reliability constraints. Our experimental results show average savings of 35–80 % in application service times and 13–15 % in energy consumption, compared to prior work.

  • Research Article
  • Cite Count Icon 2
  • 10.1109/tvlsi.2016.2594238
A Runtime Framework for Robust Application Scheduling With Adaptive Parallelism in the Dark-Silicon Era
  • Feb 1, 2017
  • IEEE Transactions on Very Large Scale Integration (VLSI) Systems
  • Nishit Kapadia + 1 more

With deeper technology scaling accompanied by a worsening power wall, an increasing proportion of chip area on a chip multiprocessor (CMP) is expected to be occupied by dark silicon. At the same time, design challenges due to process variations and soft errors in integrated circuits are projected to become even more severe. It is well known that spatial variations in process parameters introduce significant unpredictability in the performance and power profiles of CMP cores. By mapping applications onto the best set of cores, process variations can potentially be used to our advantage in the dark-silicon era. In addition, the probability of occurrence of soft errors during application execution has been found to be strongly related to the supply voltage and operating frequency values, thus necessitating reliability awareness within runtime voltage scaling schemes in contemporary CMPs. In this paper, we present a novel framework that leverages the knowledge of variations on the chip to perform runtime application mapping and dynamic voltage scaling to optimize system performance and energy, while satisfying dark-silicon power constraints of the chip as well as application-specific performance and reliability constraints. Our experimental results show average savings of 10%-71% in application service times and 13%-38% in energy consumption, compared with prior work.

  • Conference Article
  • Cite Count Icon 57
  • 10.1145/2463209.2488948
HaDeS
  • May 29, 2013
  • Yatish Turakhia + 3 more

In this paper, we propose an efficient iterative optimization based approach for architectural synthesis of dark silicon heterogeneous chip multi-processors (CMPs). The goal is to determine the optimal number of cores of each type to provision the CMP with, such that the area and power budgets are met and the application performance is maximized. We consider general-purpose multi-threaded applications with a varying degree of parallelism (DOP) that can be set at run-time, and propose an accurate analytical model to predict the execution time of such applications on heterogeneous CMPs. Our experimental results illustrate that the synthesized heterogeneous dark silicon CMPs provide between 19% to 60% performance improvements over conventional homogeneous designs for variable and fixed DOP scenarios, respectively.

  • Research Article
  • Cite Count Icon 3
  • 10.1007/s11227-015-1448-2
Leveraging dark silicon to optimize networks-on-chip topology
  • Jun 9, 2015
  • The Journal of Supercomputing
  • Mehdi Modarressi + 1 more

This paper presents a reconfigurable network-on-chip (NoC) for many-core chip multiprocessors (CMPs) in the dark silicon era, where a considerable part of high-end chips cannot be powered up due to the power and bandwidth walls. Core specialization, which trades off the cheaper silicon area with energy-efficiency, is a promising solution to the dark silicon challenge. This approach integrates a selection of many diverse application-specific cores into a single many-core chip. Each application then activates those cores that best match its processing requirements. Since active cores may not always form a contiguous active region in the chip, such a partially active many-core CMP requires some special on-chip communication support to optimize NoC parameters for the current set of active cores. In this paper, we propose a reconfigurable NoC that leverages inactive routers of a many-core chip to customize the topology for active cores. In this design, routers of the dark part of the chip are used as bypass switches that can set up virtual long links between distant active nodes in the network. Our experimental results show considerable reduction in NoC energy consumption and latency.

  • Conference Article
  • Cite Count Icon 26
  • 10.5555/2523721.2523743
A unified view of non-monotonic core selection and application steering in heterogeneous chip multiprocessors
  • Oct 7, 2013
  • Sandeep Navada + 3 more

A single-ISA heterogeneous chip multiprocessor (HCMP) is an attractive substrate to improve single-thread performance and energy efficiency in the dark silicon era. We consider HCMPs comprised of non-monotonic types where each type is performance-optimized to different instruction-level behavior and hence cannot be ranked -- different program phases achieve their highest performance on different cores. Although non-monotonic heterogeneous designs offer higher performance potential than either monotonic heterogeneous designs or homogeneous designs, steering applications to the best-performing is challenging due to performance ambiguity of types. In this paper, we present a unified view of selecting non-monotonic types at design-time and steering program phases to cores at run-time. After comprehensive evaluation, we found that with N types, the optimal HCMP for single-thread performance is comprised of an core type coupled with N-1 core types that relieve distinct resource bottlenecks in the average core. This inspires a complementary steering algorithm in which a running program is continuously diagnosed for bottlenecks on the current core. If any are observed, the program is migrated to an accelerator that relieves any of the bottlenecks and does not worsen any of them. If no accelerator satisfies this condition, then the average is selected. In our evaluation, we show that a 4-core-type HCMP improves single-thread performance up to 76% and 15% on average over a homogeneous chip multiprocessor, and our steering algorithm is able to capture most of this performance gain. Further, we show that our steering algorithm on a 4-core-type HCMP is, on average, 33% more power-efficient (BIPS3/watt) than a homogeneous chip multiprocessor.

  • Conference Article
  • Cite Count Icon 4
  • 10.1109/vlsi-sata.2016.7593059
OptMem: Dark-silicon aware low latency hybrid memory design
  • Jan 1, 2016
  • Salman Onsori + 3 more

In this article, we present a convex optimization model to design a three dimension (3D)stacked hybrid memory system to improve performance in the dark silicon era. Our convex model optimizes numbers and placement of static random access memory (SRAM) and spin-transfer torque magnetic random-access memory(STT-RAM) memories on the memory layer to exploit advantages of both technologies. Power consumption that is the main challenge in the dark silicon era is represented as a main constraint in this work and it is satisfied by the detailed optimization model in order to design a dark silicon aware 3D Chip-Multiprocessor (CMP). Experimental results show that the proposed architecture improves the energy consumption and performanceof the 3D CMPabout 25.8% and 12.9% on averagecompared to the Baseline memory design.

  • Conference Article
  • Cite Count Icon 5
  • 10.1109/ipdpsw.2012.18
Performance Benefits of Heterogeneous Computing in HPC Workloads
  • May 1, 2012
  • Victor W Lee + 2 more

Chip multi-processors (CMPs) with increasing number of processor cores are now becoming widely available. To take advantage of many-core CMPs, applications must be parallelized. However, due to the nature of algorithm / programming model, some parts of the application would remain serial. According to Amdahl's law, the speedup of a parallel application is limited by the amount of serial execution it has. For a CMP with many cores, this can be a serious limitation. To take full advantage of the increasing number of cores, one must try to reduce the execution time of the serial portion of a parallel program. However, rewriting an application takes time and often the return on the effort invested may not justify parallelizing every part of the program. Heterogeneous many-core CMP design is one possible solution to support massive parallel execution and to provide a reasonable single-thread performance. In this paper, we use a simple spreadsheet model to evaluate homogeneous and heterogeneous CMP designs using execution profiles of real HPC applications. Evaluated on 12 parallel HPC applications, we show that heterogeneous CMPs can outperform homogeneous CMPs by up to 1.35x with an average speedup of 1.06x when both the heterogeneous CMPs and homogeneous CMPs are constrained to use the same power budget. Our study found the heterogeneous CMPs can take advantage of serial portion of execution that is as little as 2% of total run time to provide performance benefit. This suggests heterogeneous computing can help mitigate the effect of not parallelizing some portions of an application due to return on investment concern on programming efforts.

  • Conference Article
  • Cite Count Icon 42
  • 10.1109/pact.2013.6618811
Jigsaw: scalable software-defined caches
  • Oct 1, 2013
  • Sandeep Navada + 3 more

A single-ISA heterogeneous chip multiprocessor (HCMP) is an attractive substrate to improve single-thread performance and energy efficiency in the dark silicon era. We consider HCMPs comprised of non-monotonic core types where each core type is performance-optimized to different instruction-level behavior and hence cannot be ranked -- different program phases achieve their highest performance on different cores. Although non-monotonic heterogeneous designs offer higher performance potential than either monotonic heterogeneous designs or homogeneous designs, steering applications to the best-performing core is challenging due to performance ambiguity of core types. In this paper, we present a unified view of selecting non-monotonic core types at design-time and steering program phases to cores at run-time. After comprehensive evaluation, we found that with N core types, the optimal HCMP for single-thread performance is comprised of an "average core" type coupled with N-1 "accelerator core" types that relieve distinct resource bottlenecks in the average core. This inspires a complementary steering algorithm in which a running program is continuously diagnosed for bottlenecks on the current core. If any are observed, the program is migrated to an accelerator core that relieves any of the bottlenecks and does not worsen any of them. If no accelerator core satisfies this condition, then the average core is selected. In our evaluation, we show that a 4-core-type HCMP improves single-thread performance up to 76% and 15% on average over a homogeneous chip multiprocessor, and our steering algorithm is able to capture most of this performance gain. Further, we show that our steering algorithm on a 4-core-type HCMP is, on average, 33% more power-efficient (BIPS3/watt) than a homogeneous chip multiprocessor.

  • Conference Article
  • Cite Count Icon 13
  • 10.1109/iiswc.2006.302744
Exploring Small-Scale and Large-Scale CMP Architectures for Commercial Java Servers
  • Oct 1, 2006
  • R Iyer + 7 more

As we enter the era of chip multiprocessor (CMP) architectures, it is important that we explore the scaling characteristics of mainstream server workloads on these platforms. In this paper, we analyze the performance of an Enterprise Java workload (SPECjbb2005) on two important classes of CMP architectures. One class of CMP platforms comprise of "small-scale" CMP (SCMP) processors with a few large out-of order cores on the die. Another class of CMP platforms comprise of "large-scale" CMP (LCMP) processors) with several small in-order cores on the die. For these classes of CMP architectures to succeed, it is important that there are sufficient resources (cache, memory and interconnect) to allow for a balanced scalable platform. In this paper, we focus on evaluating the resource scaling characteristics (cores, caches and memory) of SPECjbb2005 on these two architectures and understanding architectural trade-offs that may be required in future CMP offerings. The overall evaluation is uniquely conducted using four different methodologies (measurements on latest platforms, trace-based cache simulation, trace-based platform simulation and execution-driven emulation). Based on our findings, we summarize the architectural recommendations for future CMP server platforms (e.g. the need for large DRAM caches)

  • Conference Article
  • Cite Count Icon 1
  • 10.1109/iscas.2016.7539127
High performance 3D CMP design with stacked hybrid memory architecture in the dark silicon era using a convex optimization model
  • May 1, 2016
  • Salman Onsori + 3 more

In this article, we present a convex optimization model to design a stacked hybrid memory system to improve performance and reduce energy consumption of the chip-multiprocessor (CMP). Our convex model optimizes numbers and placement of SRAM and STT-RAM memories on the memory layer, and efficiently maps applications/threads on cores in the core layer. Power consumption that is the main challenge in the dark silicon era is represented as a power constraint in this work and it is satisfied by the detailed optimization model in order to design a dark silicon aware 3D CMP. Experimental results show that the proposed architecture considerably improves the energy-delay product (EDP) and performance of the 3D CMP compared to the Baseline memory design.

Save Icon
Up Arrow
Open/Close
Setting-up Chat
Loading Interface