K-Means-based isolation forest

Paweł Karczmarek,Adam Kiersztyn,Witold Pedrycz,Ebru Al

doi:10.1016/j.knosys.2020.105659

Abstract

The task of anomaly detection in data is one of the main challenges in data science because of the wide plethora of applications and despite a spectrum of available methods. Unfortunately, many of anomaly detection schemes are still imperfect i.e., they are not effective enough or act in a non-intuitive way or they are focused on a specific type of data. In this study, the classical method of Isolation Forest is thoroughly analyzed and augmented by bringing an innovative approach. This is k-Means-Based Isolation Forest that allows to build a search tree based on many branches in contrast to the only two considered in the original method. k-Means clustering is used to predict the number of divisions on each decision tree node. As supported through experimental studies, the proposed method works effectively for data coming from various application areas including intermodal transport and geographical, spatio-temporal data. In addition, it enables a user to intuitively determine the anomaly score for an individual record of the analyzed dataset. The advantage of the proposed method is that it is able to fit the data at the step of decision tree building. Moreover, it returns more intuitively appealing anomaly score values.

Full Text

Published Version

View

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Knowledge-Based Systems	Publication Date: Feb 17, 2020
Citations: 100	License type: cc-by-nc-nd

R Discovery Prime

K-Means-based isolation forest

Abstract

Published Version

Talk to us

Similar Papers

More From: Knowledge-Based Systems

Lead the way for us

Similar Papers

Hypersphere for Branching Node for the Family of Isolation Forest Algorithms
Jayanta Choudhury ... Piseth Ky
-
Jayanta Choudhury, et. al.Jayanta Choudhury ... Piseth Ky
01 Aug 2021
01 Aug 2021

Time and space complexity of deterministic and nondeterministic decision trees
Mikhail Moshkov
Annals of Mathematics and Artificial Intelligence | VOL. 91
Mikhail MoshkovMikhail Moshkov
09 Sep 2022
Annals of Mathematics and Artificial Intelligence | VOL. 91

Crowd-Sourcing for Data Science and Quantifiable Challenges: Optimal Contest Design
Milind Dawande ... Ganesh Janakiraman
SSRN Electronic Journal | VOL. -
Milind Dawande, et. al.Milind Dawande ... Ganesh Janakiraman
01 Jan 2020
SSRN Electronic Journal | VOL. -

Fuzzy anomaly scores for Isolation Forest
Kyoungok Kim
Applied Soft Computing | VOL. 166
Kyoungok KimKyoungok Kim
02 Sep 2024
Applied Soft Computing | VOL. 166

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

K-Means-based isolation forest

Abstract

Published Version

Talk to us

Similar Papers

More From: Knowledge-Based Systems