Monitoring and Analytics at INFN Tier-1: the next step

Fabio Viola,Enrico Fattibene,Diego Michelotto,Simone Rossi Tisbeni,Stefano Dal Pra,Lucia Morganti,Luca Dell’Agnello,Barbara Martelli,Daniele Bonacorsi,Antonio Falabella,C Doglioni,D Kim,G.A Stewart,P Jackson,W Kamleh,L Silvestris

doi:10.1051/epjconf/202024507008

Fabio Viola, Enrico Fattibene + Show 14 more

Open Access

https://doi.org/10.1051/epjconf/202024507008

Copy DOI

Journal: EPJ web of conferences	Publication Date: Jan 1, 2020
Citations: 1	License type: CC BY 4.0

Affiliation: University of Bologna

Abstract

In modern data centres an effective and efficient monitoring system is a critical asset, yet a continuous concern for administrators. Since its birth, INFN Tier-1 data centre, hosted at CNAF, has used various monitoring tools all replaced, a few years ago, by a system common to all CNAF departments (based on Sensu, Influxdb, Grafana). Given the complexity of the inter-dependencies of the several services running at the data centre and the foreseen large increase of resources in the near future, a more powerful and versatile monitoring system is needed. This new monitoring system should be able to automatically correlate log files and metrics coming from heterogeneous sources and devices (including services, hardware and infrastructure) thus providing us with a suitable framework to implement a solution for the predictive analysis of the status of the whole environment. In particular, the possibility to correlate IT infrastructure monitoring information with the logs of running applications is of great relevance in order to be able to quickly find application failure root cause. At the same time, a modern, flexible and user-friendly analytics solution is needed in order to enable users, IT engineers and IT managers to extract valuable information from the different sources of collected data in a timely fashion. In this paper, a prototype of such a system, installed at the INFN Tier-1, is described with an assessment of the state and an evaluation of the resources needed for a fully production system. Technologies adopted, amount of foreseen data, target KPIs and production design are illustrated.

Highlights

Maintenance is the process of preserving or restoring the good operating conditions of a system
INFN-CNAF’s Tier-1 site is the major Italian data centre in the Worldwide LHC Computing Grid1 (WLCG). This data centre features 40,000 CPU cores, 40 PB of disk storage, 90 PB of tape storage, and it is connected to the Italian (GARR) 2 and European (GEANT) 3 research network infrastructure with more than 200 Gbps. In such a scenario, adopting a predictive maintenance approach would allow INFN saving costs deriving from downtime or unnecessary hardware replacement
Log files may refer to system processes and services as well as log files produced by the HEP experiments running on the nodes

Summary

Introduction

Maintenance is the process of preserving or restoring the good operating conditions of a system. It is an expensive approach, since many hardware components could have a longer life and are instead replaced due to a scheduled intervention This motivates the approach known as conditionbased maintenance: the continuous monitoring of a set of measurements can highlight abnormal behaviors, allowing the administrator to perform a maintenance intervention before a fault occurs [1]. This data centre features 40,000 CPU cores, 40 PB of disk storage, 90 PB of tape storage, and it is connected to the Italian (GARR) 2 and European (GEANT) 3 research network infrastructure with more than 200 Gbps In such a scenario, adopting a predictive maintenance approach would allow INFN saving costs deriving from downtime ( providing a better service) or unnecessary hardware replacement.

The predictive maintenance infrastructure at a glance

Data sources

Software Components of the Big Data Cluster

Apache Kafka

Apache HDFS

Jupyter Hub

InfluxDB

Grafana

Infrastructure setup and deployment

Conclusion

Full Text

Published version (

Free)

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Monitoring and Analytics at INFN Tier-1: the next step

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: EPJ web of conferences

Lead the way for us

Similar Papers

Drill Bit Failure Forensics using 2D Bit Images Captured at the Rig Site
Eric Van Oort ... Zeyu Yan
-
Eric Van Oort, et. al.Eric Van Oort ... Zeyu Yan
08 Mar 2021
08 Mar 2021

Failure of plastic press release buttons in automobile seat belts
Russell F Dunn ... Brian Malone
Engineering Failure Analysis | VOL. 12
Russell F Dunn, et. al.Russell F Dunn ... Brian Malone
15 Jul 2004
Engineering Failure Analysis | VOL. 12

NHTSA warns drivers with physical disabilities against use of special steering control devices on air bag-equipped vehicles
-
Disaster Management & Response | VOL. 1
--
01 Jul 1995
Disaster Management & Response | VOL. 1

Analysis of Root Cause of Failure of a Turbo Generator Stator Winding
Dillip Kumar Puhan ... Rajat Sharma
-
Dillip Kumar Puhan, et. al.Dillip Kumar Puhan ... Rajat Sharma
03 Dec 2021
03 Dec 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Monitoring and Analytics at INFN Tier-1: the next step

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: EPJ web of conferences