Dynamic order Markov model for categorical sequence clustering

Rongbo Chen,Shengrui Wang,Haojun Sun,Jianfei Zhang,Lifei Chen

doi:10.1186/s40537-021-00547-2

Rongbo Chen, Shengrui Wang + Show 3 more

Open Access

https://doi.org/10.1186/s40537-021-00547-2

Copy DOI

Abstract

Markov models are extensively used for categorical sequence clustering and classification due to their inherent ability to capture complex chronological dependencies hidden in sequential data. Existing Markov models are based on an implicit assumption that the probability of the next state depends on the preceding context/pattern which is consist of consecutive states. This restriction hampers the models since some patterns, disrupted by noise, may be not frequent enough in a consecutive form, but frequent in a sparse form, which can not make use of the information hidden in the sequential data. A sparse pattern corresponds to a pattern in which one or some of the state(s) between the first and last one in the pattern is/are replaced by wildcard(s) that can be matched by a subset of values in the state set. In this paper, we propose a new model that generalizes the conventional Markov approach making it capable of dealing with the sparse pattern and handling the length of the sparse patterns adaptively, i.e. allowing variable length pattern with variable wildcards. The model, named Dynamic order Markov model (DOMM), allows deriving a new similarity measure between a sequence and a set of sequences/cluster. DOMM builds a sparse pattern from sub-frequent patterns that contain significant statistical information veiled by the noise. To implement DOMM, we propose a sparse pattern detector (SPD) based on the probability suffix tree (PST) capable of discovering both sparse and consecutive patterns, and then we develop a divisive clustering algorithm, named DMSC, for Dynamic order Markov model for categorical sequence clustering. Experimental results on real-world datasets demonstrate the promising performance of the proposed model.

Highlights

Categorical sequence data have grown enormously in commercial and scientific studies over the past decades
This is followed by a detailed description of the proposed model, including Dynamic Order Markov Model (DOMM), sparse pattern detector (SPD), the clustering optimizer based on Dynamic order Markov model (DOMM) and the divisive algorithm DMSC; and we describe the experimental results and analysis on the performance of DMSC
To show the performance of DMSC, three performance metrics accuracy, F1-measure and Normalized Mutual Information (NMI) obtained on the test datasets by our model and the baselines are demonstrated in Tables 2, 3 and 4, respectively

Summary

Introduction

Categorical sequence data have grown enormously in commercial and scientific studies over the past decades. To best of our knowledge, the existing sequence clustering methods are either focus on detecting the frequent consecutive patterns or sparse patterns for categorical sequence analysis, which results in overfitting or underfitting problem in terms of knowledge discovery and representation.

Results

Conclusion

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Journal of Big Data	Publication Date: Dec 1, 2021
Citations: 4	License type: open-access

R Discovery Prime

R Discovery Prime

Dynamic order Markov model for categorical sequence clustering

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Journal of Big Data

Lead the way for us

Similar Papers

Predict Gas Drainage Quantity of Mining with Dynamic Unbiased Grey Markov Model
Qing Song Li ... Shao Quan Li
Advanced Materials Research | VOL. 734-737
Qing Song Li, et. al.Qing Song Li ... Shao Quan Li
16 Aug 2013
Advanced Materials Research | VOL. 734-737

Dynamic Markov Model: Password Guessing Using Probability Adjustment Method
Xiaozhou Guo ... Huaxiang Lu
Applied Sciences | VOL. 11
Xiaozhou Guo, et. al.Xiaozhou Guo ... Huaxiang Lu
18 May 2021
Applied Sciences | VOL. 11

A hidden Markov model of electroencephalographic brain activity for advanced EEG-based brain computer interfaces
Hilmi Yanar ... Yuriy Mishchenko
-
Hilmi Yanar, et. al.Hilmi Yanar ... Yuriy Mishchenko
01 May 2016
01 May 2016

Anomaly detection based on a dynamic Markov model
Huorong Ren ... Zhiwu Li
Information Sciences | VOL. 411
Huorong Ren, et. al.Huorong Ren ... Zhiwu Li
15 May 2017
Information Sciences | VOL. 411

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Dynamic order Markov model for categorical sequence clustering

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Journal of Big Data