Disk storage management for LHCb based on Data Popularity estimator

Mikhail Hushchyn,Andrey Ustyuzhanin,Philippe Charpentier

doi:10.1088/1742-6596/664/4/042026

Abstract

This paper presents an algorithm providing recommendations for optimizing the LHCb data storage. The LHCb data storage system is a hybrid system. All datasets are kept as archives on magnetic tapes. The most popular datasets are kept on disks. The algorithm takes the dataset usage history and metadata (size, type, configuration etc.) to generate a recommendation report. This article presents how we use machine learning algorithms to predict future data popularity. Using these predictions it is possible to estimate which datasets should be removed from disk. We use regression algorithms and time series analysis to find the optimal number of replicas for datasets that are kept on disk. Based on the data popularity and the number of replicas optimization, the algorithm minimizes a loss function to find the optimal data distribution. The loss function represents all requirements for data distribution in the data storage system. We demonstrate how our algorithm helps to save disk space and to reduce waiting times for jobs using this data.

Highlights

In this module the data popularity and the predicted future usage intensities are used to estimate which datasets should be kept on disk and how many replicas they should have
The LHCb collaboration is one of the four major experiments at the Large Hadron Collider at CERN
In the results section we show a comparison of our algorithm with a simple Last Recently Used (LRU) algorithm

Summary

Introduction

In this module the data popularity and the predicted future usage intensities are used to estimate which datasets should be kept on disk and how many replicas they should have. Dataset usage history represents as time series of 104 points. The Nadaraya-Watson equation for kernel smoothing with LOO smoothing window width optimization is applied to time series of dataset usage history.

Results

Conclusion

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Journal of Physics: Conference Series	Publication Date: Dec 1, 2015
Citations: 10	License type: cc-by

R Discovery Prime

R Discovery Prime

Disk storage management for LHCb based on Data Popularity estimator

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Journal of Physics: Conference Series

Lead the way for us

Similar Papers

Distributed Data Replication and Access Optimization for LHCb Storage System - A Position Paper
Andrey Ustyuzhanin ... Philippe Charpentier
-
Andrey Ustyuzhanin, et. al.Andrey Ustyuzhanin ... Philippe Charpentier
01 Jan 2015
01 Jan 2015

Utilizing cloud storage architecture for long-pulse fusion experiment data storage
Ming Zhang ... Kexun Yu
Fusion Engineering and Design | VOL. 112
Ming Zhang, et. al.Ming Zhang ... Kexun Yu
22 Feb 2016
Fusion Engineering and Design | VOL. 112

Load Balancing Cloud Storage Data Distribution Strategy of Internet of Things Terminal Nodes considering Access Cost.
Jiansheng Wu ... Akshi Kumar
Computational Intelligence and Neuroscience | VOL. 2022
Jiansheng Wu, et. al.Jiansheng Wu ... Akshi Kumar
24 Jan 2022
Computational Intelligence and Neuroscience | VOL. 2022

Analysis of search and replication in unstructured peer-to-peer networks
Saurabh Tewari ... Leonard Kleinrock
ACM SIGMETRICS Performance Evaluation Review | VOL. 33
Saurabh Tewari, et. al.Saurabh Tewari ... Leonard Kleinrock
06 Jun 2005
ACM SIGMETRICS Performance Evaluation Review | VOL. 33

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Disk storage management for LHCb based on Data Popularity estimator

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Journal of Physics: Conference Series