A Data Layout and Fast Failure Recovery Scheme for Distributed Storage Systems with Mixed Erasure Codes

Liangliang Xu,Yinlong Xu,Zhipeng Li,Cheng Li,Min Lyu

doi:10.1109/tc.2021.3105882

Abstract

Erasure coding becomes increasingly popular in distributed storage systems (DSSes) for providing high reliability with low storage overhead. However, traditional random data placement induces massive cross-rack traffic and severely imbalanced load during failure recovery, which degrades the recovery performance significantly. In addition, various erasure codes coexisting in a DSS exacerbates the above problems. In this paper, we propose PDL, a PBD-based Data Layout, to optimize failure recovery performance in DSSes. PDL is constructed based on Pairwise Balanced Design, a combinatorial design scheme with uniform mathematical properties, and thus presents a uniform data layout for mixed erasure codes. Then we propose rPDL, a failure recovery scheme based on PDL. rPDL reduces cross-rack traffic effectively and provides nearly balanced cross-rack traffic distribution by uniformly choosing replacement nodes and retrieving determined available blocks to recover the lost blocks. We implemented PDL and rPDL in Hadoop 3.1.1. Compared with the existing data layout and recovery scheme in HDFS, experimental results show that rPDL achieves much higher recovery throughput, 6.27x for single-node failures, 5.14x for multi-node failures and 1.48x for single-rack failures, respectively. It also reduces degraded read latency by 62.83%, and provides evidently better support to front-end applications in case of component failures.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

A Data Layout and Fast Failure Recovery Scheme for Distributed Storage Systems with Mixed Erasure Codes

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Computers

Lead the way for us

Journal: IEEE Transactions on Computers	Publication Date: Jan 1, 2021
Citations: 1

Similar Papers

PDL: A Data Layout towards Fast Failure Recovery for Erasure-coded Distributed Storage Systems
Liangliang Xu ... Zhipeng Li
-
Liangliang Xu, et. al.Liangliang Xu ... Zhipeng Li
01 Jul 2020
01 Jul 2020

D3: Deterministic Data Distribution for Efficient Data Reconstruction in Erasure-Coded Distributed Storage Systems
Zhipeng Li ... Liangliang Xu
-
Zhipeng Li, et. al.Zhipeng Li ... Liangliang Xu
01 May 2019
01 May 2019

Zebra: Demand-aware erasure coding for distributed storage systems
Jun Li ... Baochun Li
-
Jun Li, et. al.Jun Li ... Baochun Li
01 Jun 2016
01 Jun 2016

Observation Likelihood Model Design and Failure Recovery Scheme Toward Reliable Localization of Mobile Robots
Chang-Bae Moon ... Woojin Chung
International Journal of Advanced Robotic Systems | VOL. 7
Chang-Bae Moon, et. al.Chang-Bae Moon ... Woojin Chung
01 Jan 2009
International Journal of Advanced Robotic Systems | VOL. 7

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Data Layout and Fast Failure Recovery Scheme for Distributed Storage Systems with Mixed Erasure Codes

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Computers