Anomaly Detection Methods for Categorical Data

Ayman Taha,Ali S Hadi

doi:10.1145/3312739

Abstract

Anomaly detection has numerous applications in diverse fields. For example, it has been widely used for discovering network intrusions and malicious events. It has also been used in numerous other applications such as identifying medical malpractice or credit fraud. Detection of anomalies in quantitative data has received a considerable attention in the literature and has a venerable history. By contrast, and despite the widespread availability use of categorical data in practice, anomaly detection in categorical data has received relatively little attention as compared to quantitative data. This is because detection of anomalies in categorical data is a challenging problem. Some anomaly detection techniques depend on identifying a representative pattern then measuring distances between objects and this pattern. Objects that are far from this pattern are declared as anomalies. However, identifying patterns and measuring distances are not easy in categorical data compared with quantitative data. Fortunately, several papers focussing on the detection of anomalies in categorical data have been published in the recent literature. In this article, we provide a comprehensive review of the research on the anomaly detection problem in categorical data. Previous review articles focus on either the statistics literature or the machine learning and computer science literature. This review article combines both literatures. We review 36 methods for the detection of anomalies in categorical data in both literatures and classify them into 12 different categories based on the conceptual definition of anomalies they use. For each approach, we survey anomaly detection methods, and then show the similarities and differences among them. We emphasize two important issues, the number of parameters each method requires and its time complexity. The first issue is critical, because the performance of these methods are sensitive to the choice of these parameters. The time complexity is also very important in real applications especially in big data applications. We report the time complexity if it is reported by the authors of the methods. If it is not, then we derive it ourselves and report it in this article. In addition, we discuss the common problems and the future directions of the anomaly detection in categorical data.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Anomaly Detection Methods for Categorical Data

Abstract

Talk to us

Similar Papers

More From: ACM Computing Surveys

Lead the way for us

Journal: ACM Computing Surveys	Publication Date: May 30, 2019
Citations: 68

Similar Papers

Automated anomaly detection for categorical data by repurposing a form filling recommender system
Hichem Belgacem ... Xiaochen Li
Journal of Data and Information Quality | VOL. -
Hichem Belgacem, et. al.Hichem Belgacem ... Xiaochen Li
16 Sep 2024
Journal of Data and Information Quality | VOL. -

SHARD: A Framework for Sequential, Hierarchical Anomaly Ranking and Detection
Jason Robinson ... Allison Candido
-
Jason Robinson, et. al.Jason Robinson ... Allison Candido
01 Jan 2012
01 Jan 2012

Hybrid data-driven outlier detection based on neighborhood information entropy and its developmental measures
Zhong Yuan ... Shan Feng
Expert Systems with Applications | VOL. 112
Zhong Yuan, et. al.Zhong Yuan ... Shan Feng
07 Jun 2018
Expert Systems with Applications | VOL. 112

Clustered Hierarchical Anomaly and Outlier Detection Algorithms
Najib Ishaq ... Noah M Daniels
-
Najib Ishaq, et. al.Najib Ishaq ... Noah M Daniels
15 Dec 2021
15 Dec 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Anomaly Detection Methods for Categorical Data

Abstract

Talk to us

Similar Papers

More From: ACM Computing Surveys