Declarative Resilience

Hamza Omar,Omer Khan,Qingchuan Shi,Masab Ahmad,Halit Dogan

doi:10.1145/3210559

Abstract

To protect multicores from soft-error perturbations, research has explored various resiliency schemes that provide high soft-error coverage. However, these schemes incur high performance and energy overheads. We observe that not all soft-error perturbations affect program correctness, and some soft-errors only affect program accuracy, i.e., the program completes with certain acceptable deviations from error free outcome. Thus, it is practical to improve processor efficiency by trading off resiliency overheads with program accuracy. This article proposes the idea of declarative resilience that selectively applies strong resiliency schemes for code regions that are crucial for program correctness (crucial code) and lightweight resiliency for code regions that are susceptible to program accuracy deviations as a result of soft-errors (non-crucial code). At the application level, crucial and non-crucial code is identified based on its impact on the program outcome. A cross-layer architecture enables efficient resilience along with holistic soft-error coverage. Only program accuracy is compromised in the worst-case scenario of a soft-error strike during non-crucial code execution. For a set of machine-learning and graph analytic benchmarks, declarative resilience reduces performance overhead over a state-of-the-art system that applies strong resiliency for all program code regions from ∼ 1.43× to ∼ 1.2×.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Declarative Resilience

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Embedded Computing Systems

Lead the way for us

Journal: ACM Transactions on Embedded Computing Systems	Publication Date: Jul 24, 2018
Citations: 7

Similar Papers

A Cross-Layer Multicore Architecture to Tradeoff Program Accuracy and Resilience Overheads
Qingchuan Shi ... Omer Khan
IEEE Computer Architecture Letters | VOL. 14
Qingchuan Shi, et. al.Qingchuan Shi ... Omer Khan
01 Jul 2015
IEEE Computer Architecture Letters | VOL. 14

FaultHound
Nitin ... T N Vijaykumar
-
Nitin, et. al. Nitin ... T N Vijaykumar
13 Jun 2015
13 Jun 2015

FaultHound
Nitin ... Irith Pomeranz
ACM SIGARCH Computer Architecture News | VOL. 43
Nitin, et. al. Nitin ... Irith Pomeranz
13 Jun 2015
ACM SIGARCH Computer Architecture News | VOL. 43

Power and thermal characterization of POWER6 system
Victor Jiménez ... Eren Kursun
-
Victor Jiménez, et. al.Victor Jiménez ... Eren Kursun
11 Sep 2010
11 Sep 2010

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Declarative Resilience

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Embedded Computing Systems