Entropy-based outlier detection using spark

Guilan Feng,Zhengnan Li,Wengang Zhou,Shi Dong

doi:10.1007/s10586-019-02932-2

Abstract

The k-nearest neighbors outlier detection is a simple yet effective widely renowned method in data mining. The actual application of this model in the big data domain is not feasible due to time and memory restrictions. Several distributed alternatives based on MapReduce have been proposed to enable this method to handle large-scale data. However, their performance can be further improved with new designs that fit with newly arising technologies. Furthermore, it gives to each attribute the same importance to outlier. There are several approaches to enhance its precision, with the entropy-based outlier detection being among the most successful ones. Entropy-based outlier detection computes attribute entropy of the data set to weighted distance formula for the outlier detection. Apart from the existing the k-nearest neighbors outlier detection to handle big datasets, there is not an entropy-based outlier detection to manage that volume of data. In this paper, we propose an entropy-based outlier detection based on Spark. It presents three separately stages. The first stage computes attribute entropy. The second stage finds the k nearest neighbors and calculates the degrees of outliers using the attribute entropy computed previously. The third stage ranks each point on the degrees of outliers and declares the top n points in this ranking to be outliers. Extensive experimental results show the advantages of the proposed method. This algorithm can improve the outlier detection precision, reduce the runtime and realize the effective large scale dataset outlier detection.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Entropy-based outlier detection using spark

Abstract

Talk to us

Similar Papers

More From: Cluster Computing

Lead the way for us

Journal: Cluster Computing	Publication Date: Apr 16, 2019
Citations: 4

Similar Papers

Introduction to Data Mining

Scalable Computing Practice and Experience | VOL. 9

01 Jan 2008
Scalable Computing Practice and Experience | VOL. 9

Harmonisation de la prise en charge respiratoire des patients atteints de SLA en France
J Gonzalez-Bermejo ... T Perez
Revue des Maladies Respiratoires | VOL. 22
J Gonzalez-Bermejo, et. al.J Gonzalez-Bermejo ... T Perez
01 Feb 2005
Revue des Maladies Respiratoires | VOL. 22

KNN-IS: An Iterative Spark-based design of the k-Nearest Neighbors classifier for big data
Jesus Maillo ... Francisco Herrera
Knowledge-Based Systems | VOL. 117
Jesus Maillo, et. al.Jesus Maillo ... Francisco Herrera
14 Jun 2016
Knowledge-Based Systems | VOL. 117

Neighborhood outlier detection
Yumin Chen ... Hongyun Zhang
Expert Systems with Applications | VOL. 37
Yumin Chen, et. al.Yumin Chen ... Hongyun Zhang
06 Jul 2010
Expert Systems with Applications | VOL. 37

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Entropy-based outlier detection using spark

Abstract

Talk to us

Similar Papers

More From: Cluster Computing