Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

The BOSS is concerned with time series classification in the presence of noise

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Similarity search is one of the most important and probably best studied methods for data mining. In the context of time series analysis it reaches its limits when it comes to mining raw datasets. The raw time series data may be recorded at variable lengths, be noisy, or are composed of repetitive substructures. These build a foundation for state of the art search algorithms. However, noise has been paid surprisingly little attention to and is assumed to be filtered as part of a preprocessing step carried out by a human. Our Bag-of-SFA-Symbols (BOSS) model combines the extraction of substructures with the tolerance to extraneous and erroneous data using a noise reducing representation of the time series. We show that our BOSS ensemble classifier improves the best published classification accuracies in diverse application areas and on the official UCR classification benchmark datasets by a large margin.

Similar Papers
  • Research Article
  • Cite Count Icon 141
  • 10.1016/j.neucom.2019.06.032
A deep learning framework for time series classification using Relative Position Matrix and Convolutional Neural Network
  • Jun 13, 2019
  • Neurocomputing
  • Wei Chen + 1 more

A deep learning framework for time series classification using Relative Position Matrix and Convolutional Neural Network

  • Conference Article
  • Cite Count Icon 20
  • 10.18293/seke2016-067
Time Series Classification with Discrete Wavelet Transformed Data: Insights from an Empirical Study
  • Jul 1, 2016
  • Proceedings/Proceedings of the ... International Conference on Software Engineering and Knowledge Engineering
  • Daoyuan Li + 3 more

Time series mining has become essential for extracting knowledge from the abundant data that flows out from many application domains. To overcome storage and processing challenges in time series mining, compression techniques are being used. In this paper, we investigate the loss/gain of performance of time series classification approaches when fed with lossy-compressed data. This empirical study is essential for reassuring practitioners, but also for providing more insights on how compression techniques can even be effective in reducing noise in time series data. From a knowledge engineering perspective, we show that time series may be compressed by 90% using discrete wavelet transforms and still achieve remarkable classification accuracy, and that residual details left by popular wavelet compression techniques can sometimes even help achieve higher classification accuracy than the raw time series data, as they better capture essential local features.

  • Research Article
  • Cite Count Icon 23
  • 10.1142/s0218194016400088
Time Series Classification with Discrete Wavelet Transformed Data
  • Nov 1, 2016
  • International Journal of Software Engineering and Knowledge Engineering
  • Daoyuan Li + 3 more

Time series mining has become essential for extracting knowledge from the abundant data that flows out from many application domains. To overcome storage and processing challenges in time series mining, compression techniques are being used. In this paper, we investigate the loss/gain of performance of time series classification approaches when fed with lossy-compressed data. This extended empirical study is essential for reassuring practitioners, but also for providing more insights on how compression techniques can even be effective in smoothing and reducing noise in time series data. From a knowledge engineering perspective, we show that time series may be compressed by 90% using discrete wavelet transforms and still achieve remarkable classification accuracy, and that residual details left by popular wavelet compression techniques can sometimes even help to achieve higher classification accuracy than the raw time series data, as they better capture essential local features.

  • Conference Article
  • 10.62036/isd.2025.36
Classification of SAX-like Time Series Bitmaps with Siamese Convolutional Neural Networks
  • Nov 17, 2025
  • Proceedings of the International Conference on Information Systems Development
  • Mariusz Wrzesien + 2 more

This research paper concerns the classification of multiclass univariate time series. It extends the existing time series imaging method and describes the process of obtaining new transformations of this sort. The presented techniques leverage the well-known Symbolic Aggregate Approximation representation, which transforms time series from numerical to symbolic domain. The obtained symbolic approximations are later turned into images. After raw time series data is transformed into two-dimensional grayscale bitmaps, these bitmaps are used as input for two alternative deep learning classification approaches. Our study focuses on comparing regular Convolutional Neural Networks and Siamese Neural Networks as time series classifiers. Experimental studies and comparative analyses were performed on well-known, publicly available datasets.

  • Research Article
  • Cite Count Icon 11
  • 10.1142/s021800142050010x
Feature-Based Online Representation Algorithm for Streaming Time Series Similarity Search
  • Sep 5, 2019
  • International Journal of Pattern Recognition and Artificial Intelligence
  • Peng Zhan + 5 more

With the rapid development of information technology, we have already access to the era of big data. Time series is a sequence of data points associated with numerical values and successive timestamps. Time series not only has the traditional big data features, but also can be continuously generated in a high speed. Therefore, it is very time- and resource-consuming to directly apply the traditional time series similarity search methods on the raw time series data. In this paper, we propose a novel online segmenting algorithm for streaming time series, which has a relatively high performance on feature representation and similarity search. Extensive experimental results on different typical time series datasets have demonstrated the superiority of our method.

  • Research Article
  • Cite Count Icon 19
  • 10.3233/jifs-171393
Combining raw and normalized data in multivariate time series classification with dynamic time warping
  • Jan 12, 2018
  • Journal of Intelligent & Fuzzy Systems
  • Maciej Łuczak

Data normalization is one of the most common processing methods applied to raw data before its subsequent use in data mining algorithms, classification, or clustering methods. Many procedures, particularly those that use any statistical analysis, require that data be normalized in one way or another. In the case of time series a standard method of processing raw data is z-normalization of each time series instance in the data set. For multivariate (multidimensional) time series we z-normalize each dimension (variable) individually. Although normalization brings a lot of advantages, it is easy to find examples of data sets where normalization destroys information contained in the raw data. In this paper we demonstrate, that for multivariate time series (MTS) both raw and normalized components give some information about the data and the best way of mining it is a combination of them. We focus here on multidimensional time series and their classification using the nearest neighbor method with the dynamic time warping (DTW) distance measure. We construct a parametric distance measure that is a combination of DTW on raw and z-normalized time series data. It turns out that the combined distance measure carries more information about the data than the two distance components separately. By determining an individual parameter for each data set it is possible to obtain a lower classification error than the errors of both component distance measures. We perform experiments on real data sets from many fields of science and technology. The advantage of the combined approach is confirmed by graphical and statistical comparisons.

  • Research Article
  • Cite Count Icon 14
  • 10.1109/msp.2022.3155955
Post Hoc Explainability for Time Series Classification: Toward a signal processing perspective
  • Jul 1, 2022
  • IEEE Signal Processing Magazine
  • Rami Mochaourab + 4 more

Time series data correspond to observations of phenomena that are recorded over time <xref ref-type="bibr" rid="ref1" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">[1]</xref> . Such data are encountered regularly in a wide range of applications, such as speech and music recognition, monitoring health and medical diagnosis, financial analysis, motion tracking, and shape identification, to name a few. With such a diversity of applications and the large variations in their characteristics, time series classification is a complex and challenging task. One of the fundamental steps in the design of time series classifiers is that of defining or constructing the <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">discriminant features</i> that help differentiate between classes. This is typically achieved by designing novel <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">representation techniques</i> <xref ref-type="bibr" rid="ref2" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">[2]</xref> that transform the raw time series data to a new data domain, where subsequently a classifier is trained on the transformed data, such as one-nearest neighbors <xref ref-type="bibr" rid="ref3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">[3]</xref> or random forests <xref ref-type="bibr" rid="ref4" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">[4]</xref> . In recent time series classification approaches, deep neural network models have been employed that are able to jointly learn a representation of time series and perform classification <xref ref-type="bibr" rid="ref5" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">[5]</xref> . In many of these sophisticated approaches, the discriminant features tend to be complicated to analyze and interpret, given the high degree of nonlinearity.

  • Conference Article
  • Cite Count Icon 3
  • 10.1109/biocas.2016.7833766
Incomplete electrocardiogram time series prediction
  • Oct 1, 2016
  • Weiwei Shi + 6 more

The prevalence of Big Data has led to the operation practice based on time series data from multiple sources in many practical applications. The prediction analysis of time series, a fundamental objective of time series data crunch, is an integral part for planning and decision making. However, the prediction analysis based on raw time series data is hardly satisfactory as missing values are usually involved in raw data samples, even if a collection of the existing regression models are available to handle the complete time series prediction problems. To achieve higher prediction precision with incomplete raw time series data, in this paper, we describe a new framework called ISM (Incomplete time series prediction based on Selective tensor modeling and Multi-kernel learning). ISM is composed of three steps. First, multi-source time series are fused and then a selective tensor is constructed from K most relevant raw data sets. Second, the selective tensor is further factorized by ISM with the sparsity constraint to extract the common latent factors across multiple sources. Finally, these factors serve as the training features of multi-kernel learning, which is an effective approach to build multi-source regression models. Extensive experiments on the electrocardiogram data set demonstrate that the proposed framework ISM achieves a superior performance of time series prediction with missing data.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 2
  • 10.3390/ijgi12040179
Querying Similar Multi-Dimensional Time Series with a Spatial Database
  • Apr 21, 2023
  • ISPRS International Journal of Geo-Information
  • Zheren Liu + 2 more

Similar time series search is one of the most important time series mining tasks in our daily life. As recent advances in sensor technologies accumulate abundant multi-dimensional time series data associated with multivariate quantities, it becomes a privilege to adapt similar time series searches for large-scale and multi-dimensional time series data. However, traditional similar time series search methods are mainly designed for one-dimensional time series, while advanced methods applicable for multi-dimensional time series data are largely immature and, more importantly, are not friendly to users from the domain of geography. As an alternative, we propose a novel method to search similar multi-dimensional time series with spatial databases. Compared with traditional methods that often conduct the similarity search based on features of the raw time series data sequence, the proposed method stores multi-dimensional time series as spatial objects in a spatial database, and then searches similar time series based on their spatial features. To demonstrate the validity of the proposed method, we analyzed the correlation between temporal features of the raw time series and spatial features of their corresponding spatial objects theoretically and empirically. Results indicate that the proposed method can not only support similar multi-dimensional time series searches but also markedly improve its efficiency under many specific scenarios. We believe that such a new paradigm will shed further light on the similarity search in large-scale multi-dimensional time series data, and will lower the barrier for users familiar with spatial databases to conduct complex time series mining tasks.

  • Research Article
  • Cite Count Icon 93
  • 10.1109/tcyb.2018.2789422
Multiobjective Learning in the Model Space for Time Series Classification.
  • Jan 22, 2018
  • IEEE Transactions on Cybernetics
  • Zhichen Gong + 3 more

A well-defined distance is critical for the performance of time series classification. Existing distance measurements can be categorized into two branches. One is to utilize handmade features for calculating distance, e.g., dynamic time warping, which is limited to exploiting the dynamic information of time series. The other methods make use of the dynamic information by approximating the time series with a generative model, e.g., Fisher kernel. However, previous distance measurements for time series seldom exploit the label information, which is helpful for classification by distance metric learning. In order to attain the benefits of the dynamic information of time series and the label information simultaneously, this paper proposes a multiobjective learning algorithm for both time series approximation and classification, termed multiobjective model-metric (MOMM) learning. In MOMM, a recurrent network is exploited as the temporal filter, based on which, a generative model is learned for each time series as a representation of that series. The models span a non-Euclidean space, where the label information is utilized to learn the distance metric. The distance between time series is then calculated as the model distance weighted by the learned metric. The network size is also optimized to learn parsimonious representations. MOMM simultaneously optimizes the data representation, the time series model separation, and the network size. The experiments show that MOMM achieves not only superior overall performance on uni/multivariate time series classification but also promising time series prediction performance.

  • Conference Article
  • Cite Count Icon 2
  • 10.1109/indin41052.2019.8972120
A Hierarchical Storage System for Industrial Time-Series Data
  • Jul 1, 2019
  • Kevin Villalobos + 5 more

The increasing interest among manufacturers in monitoring and analyzing industrial systems is generating a problem related to the considerable costs associated with the storage of captured data. This paper presents a three-level hierarchical architecture for time-series data storage on cloud environments that helps to decrease those costs. In the first level, new raw time-series data is stored for a short-period of time (e.g., one day) on electronic non-volatile storage such as solid-state drives (SSDs) that provide fast access for real time visualization of the latest data. In the second level, recent time series are stored for a medium-period of time (e.g., one week) on magnetic hard disk drives (HDDs) that are lower-cost devices with slower data transfer speed. In the third level, a reduced representation of the time series obtained by applying time-series reduction techniques are stored in HDDs, for a longer period of time (e.g., one year). Dealing with those reduced representations, data storage and transmission costs can be decreased, without limiting the future use of the data in different processes.The architecture has been implemented by using the top Database Management System from three different categories: Wide column store, Time series DBMS and Graph DBMS. It has been tested using industrial time series coming from a real manufacturing environment, and with three different types of queries proposed by domain experts. The performance results regarding storage space, storage costs and query time processing are shown on the paper.

  • Book Chapter
  • Cite Count Icon 16
  • 10.1007/978-3-031-24378-3_4
Fast Time Series Classification with Random Symbolic Subsequences
  • Jan 1, 2023
  • Thach Le Nguyen + 1 more

Symbolic representations of time series have proven to be effective for time series classification, with many recent approaches including BOSS, WEASEL, and MrSEQL. These classifiers use various elaborate methods to select discriminative features from symbolic representations of time series. As a result, although they have competitive results regarding accuracy, their classification models are relatively expensive to train. Most if not all of these approaches have missed an important research question: are these elaborate feature selection methods actually necessary? ROCKET, a state-of-the-art time series classifier, outperforms all of them without utilizing any feature selection techniques. In this paper, we answer this question by contrasting these classifiers with a very simple method, named MrSQM. This method samples random subsequences from symbolic representations of time series. Our experiments on 112 datasets of the UEA/UCR benchmark demonstrate that MrSQM can quickly extract useful features and learn accurate classifiers with the logistic regression algorithm. MrSQM completes training and prediction on 112 datasets in 1.5 h for an accuracy comparable to existing efficient state-of-the-art methods, e.g., MrSEQL (10 h) and ROCKET (2.5 h). Furthermore, MrSQM enables the user to trade-off accuracy and speed by controlling the type and number of symbolic representations, thus further reducing the total runtime to 20 min for a similar level of accuracy. With these results, we show that random subsequences extracted from symbolic transformations can be as effective as the more sophisticated and expensive feature selection methods proposed in previous works. We propose MrSQM as a strong baseline for future research in time series classification, especially for approaches based on symbolic representations of time series.

  • Book Chapter
  • Cite Count Icon 14
  • 10.1007/978-3-319-08979-9_18
Towards Time Series Classification without Human Preprocessing
  • Jan 1, 2014
  • Patrick Schäfer

Similarity search is a core functionality in many data mining algorithms. Over the past decade these algorithms were designed to mostly work with human assistance to extract characteristic, aligned patterns of equal length and scaling. Human assistance is not cost-effective. We propose our shotgun distance similarity metric that extracts, scales, and aligns segments from a query to a sample time series. This simplifies the classification of time series as produced by sensors. A time series is classified based on its segments at varying lengths as part of our shotgun ensemble classifier. It improves the best published accuracies on case studies in the context of bioacoustics, human motion detection, spectrographs or personalized medicine. Finally, it performs better than state of the art on the official UCR classification benchmark.

  • Conference Article
  • Cite Count Icon 11
  • 10.1109/icdim.2012.6360087
Time Series Classification Method Based on Longest Common Subsequence and Textual Approximation
  • Aug 1, 2012
  • Abdulla-Al-Maruf + 2 more

Many symbolic representations of time series have been proposed by researchers over past decades. However, it is still not enough to classify time series with high accuracy in such applications as ubiquitous systems or sensor systems. In this paper, we propose a new symbolic representation of time series called l-TAX to increase the accuracy of time series classification. A time series can be represented by term sequences in l-TAX. l-TAX is based on a document like symbolic representation of time series called TAX. We use longest common subsequence as our distance measure between textually approximated time series. During time series classification, consideration of symbol sequences increases the accuracy significantly. In our evaluation, we have demonstrated that l-TAX is effective for classification as well as searching time series data set.

  • Research Article
  • Cite Count Icon 15
  • 10.3390/s24196402
Extraction of Features for Time Series Classification Using Noise Injection
  • Oct 2, 2024
  • Sensors (Basel, Switzerland)
  • Gyu Il Kim + 1 more

Time series data often display complex, time-varying patterns, which pose significant challenges for effective classification due to data variability, noise, and imbalance. Traditional time series classification techniques frequently fall short in addressing these issues, leading to reduced generalization performance. Therefore, there is a need for innovative methodologies to enhance data diversity and quality. In this paper, we introduce a method for the extraction of features for time series classification using noise injection to address these challenges. By employing noise injection techniques for data augmentation, we enhance the diversity of the training data. Utilizing digital signal processing (DSP), we extract key frequency features from time series data through sampling, quantization, and Fourier transformation. This process enhances the quality of the training data, thereby maximizing the model’s generalization performance. We demonstrate the superiority of our proposed method by comparing it with existing time series classification models. Additionally, we validate the effectiveness of our approach through various experimental results, confirming that data augmentation and DSP techniques are potent tools in time series data classification. Ultimately, this research presents a robust methodology for time series data analysis and classification, with potential applications across a broad spectrum of data analysis problems.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant