An Evaluation of Low Overhead Time Series Preprocessing Techniques for Downstream Machine Learning

Matthew L Weiss,Andrew Prout,Michael Jones,Charles Yee,Andrew Bowne,Siddharth Samsi,Vijay Gadepally,Lindsey Mcevoy,Joseph Mcdonald,Daniel Edelman,David Bestor

doi:10.1109/hpec55821.2022.9926406

Abstract

In this paper we address the application of pre-processing techniques to multi-channel time series data with varying lengths, which we refer to as the alignment problem, for downstream machine learning. The misalignment of multi-channel time series data may occur for a variety of reasons, such as missing data, varying sampling rates, or inconsistent collection times. We consider multi-channel time series data collected from the MIT SuperCloud High Performance Computing (HPC) center, where different job start times and varying run times of HPC jobs result in misaligned data. This misalignment makes it challenging to build AI/ML approaches for tasks such as compute workload classification. Building on previous supervised classification work with the MIT SuperCloud Dataset, we address the alignment problem via three broad, low overhead approaches: sampling a fixed subset from a full time series, performing summary statistics on a full time series, and sampling a subset of coefficients from time series mapped to the frequency domain. Our best performing models achieve a classification accuracy greater than 95%, outperforming previous approaches to multi-channel time series classification with the MIT SuperCloud Dataset by 5 %. These results indicate our low overhead approaches to solving the alignment problem, in conjunction with standard machine learning techniques, are able to achieve high levels of classification accuracy, and serve as a baseline for future approaches to addressing the alignment problem, such as kernel methods.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

An Evaluation of Low Overhead Time Series Preprocessing Techniques for Downstream Machine Learning

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

(Invited) Diffusion Map Analysis of Multi-Channel Time Series Data
Takashi Nakamura ... Scott Hagan
ECS Meeting Abstracts | VOL. MA2018-03
Takashi Nakamura, et. al.Takashi Nakamura ... Scott Hagan
13 Jul 2018
ECS Meeting Abstracts | VOL. MA2018-03

Recreational sector is the dominant source of fishing mortality for oceanic fishes in the Southeast United States Atlantic Ocean
Kyle W Shertzer ... Joseph Kevin Craig
Fisheries Management and Ecology | VOL. 26
Kyle W Shertzer, et. al.Kyle W Shertzer ... Joseph Kevin Craig
05 Jul 2019
Fisheries Management and Ecology | VOL. 26

Time-series aggregation for synthesis problems by bounding error in the objective function
Björn Bahl ... André Bardow
Energy | VOL. 135
Björn Bahl, et. al.Björn Bahl ... André Bardow
18 Jun 2017
Energy | VOL. 135

Chaos, dynamical structure, and climate variability
H Bruce Stewart
-
H Bruce StewartH Bruce Stewart
01 Jan 1996
01 Jan 1996

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

An Evaluation of Low Overhead Time Series Preprocessing Techniques for Downstream Machine Learning

Abstract

Talk to us

Similar Papers