Practical online failure prediction for Blue Gene/P: Period-based vs event-driven

Li Yu,Zhiling Lan,Ziming Zheng,Susan Coghlan

doi:10.1109/dsnw.2011.5958823

Abstract

To facilitate proactive fault management in large-scale systems such as IBM Blue Gene/P, online failure prediction is of paramount importance. While many techniques have been presented for online failure prediction, questions arise regarding two commonly used approaches: period-based and event-driven. Which one has better accuracy? What is the best observation window (i.e., the time interval used to collect evidence before making a prediction)? How does the lead time (i.e., the time interval from the prediction to the failure occurrence) impact prediction arruracy? To answer these questions, we analyze and compare period-based and event-driven prediction approaches via a Bayesian prediction model. We evaluate these prediction approaches, under a variety of testing parameters, by means of RAS logs collected from a production supercomputer at Argonne National Laboratory. Experimental results show that the period-based Bayesian model and the event-driven Bayesian model can achieve up to 65.0% and 83.8% prediction accuracy, respectively. Furthermore, our sensitivity study indicates that the event-driven approach seems more suitable for proactive fault management in large-scale systems like Blue Gene/P.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Practical online failure prediction for Blue Gene/P: Period-based vs event-driven

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Explore unlabeled big data learning to online failure prediction in safety-aware cloud environment
Jia Zhao ... Ming Hu
Journal of Parallel and Distributed Computing | VOL. 153
Jia Zhao, et. al.Jia Zhao ... Ming Hu
17 Mar 2021
Journal of Parallel and Distributed Computing | VOL. 153

A survey of online failure prediction methods
Felix Salfner ... Maren Lenk
ACM Computing Surveys | VOL. 42
Felix Salfner, et. al.Felix Salfner ... Maren Lenk
01 Mar 2010
ACM Computing Surveys | VOL. 42

Prediction models for clustered data with informative priors for the random effects: a simulation study
Haifang Ni ... Irene Klugkist
BMC Medical Research Methodology | VOL. 18
Haifang Ni, et. al.Haifang Ni ... Irene Klugkist
06 Aug 2018
BMC Medical Research Methodology | VOL. 18

A Bayesian performance prediction model for mathematics education: A prototypical approach for effective group composition
Rahel Bekele ... Maggie Mcpherson
British Journal of Educational Technology | VOL. 42
Rahel Bekele, et. al.Rahel Bekele ... Maggie Mcpherson
06 Apr 2011
British Journal of Educational Technology | VOL. 42

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Practical online failure prediction for Blue Gene/P: Period-based vs event-driven

Abstract

Talk to us

Similar Papers