Big data: an optimized approach for cluster initialization

Marina Gul,M Abdul Rehman

doi:10.1186/s40537-023-00798-1

Abstract

The k-means, one of the most widely used clustering algorithm, is not only faster in computation but also produces comparatively better clusters. However, it has two major downsides, first it is sensitive to initialize k value and secondly, especially for larger datasets, the number of iterations could be very large, making it computationally hard. In order to address these issues, we proposed a scalable and cost-effective algorithm, called R-k-means, which provides an optimized solution for better clustering large scale high-dimensional datasets. The algorithm first selects O(R) initial points then reselect O(l) better initial points, using distance probability from dataset. These points are then again clustered into k initial points. An empirical study in a controlled environment was conducted using both simulated and real datasets. Experimental results showed that the proposed approach outperformed as compared to the previous approaches when the size of data increases with increasing number of dimensions.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Journal of Big Data	Publication Date: Jul 20, 2023
Citations: 5	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Big data: an optimized approach for cluster initialization

Abstract

Talk to us

Similar Papers

More From: Journal of Big Data

Lead the way for us

Similar Papers

An iterative initial-points refinement algorithm for categorical data clustering
Ying Sun ... Zhengxin Chen
Pattern Recognition Letters | VOL. 23
Ying Sun, et. al.Ying Sun ... Zhengxin Chen
06 Dec 2001
Pattern Recognition Letters | VOL. 23

New diagonal bundle method for clustering problems in large data sets
Napsu Karmitsa ... Sona Taheri
European Journal of Operational Research | VOL. 263
Napsu Karmitsa, et. al.Napsu Karmitsa ... Sona Taheri
10 Jun 2017
European Journal of Operational Research | VOL. 263

Density Based Initial Center Optimization Algorithm
Shengli Sun ... Zhigao Zheng
-
Shengli Sun, et. al.Shengli Sun ... Zhigao Zheng
01 Jan 2013
01 Jan 2013

A shrinking synchronization clustering algorithm based on a linear weighted Vicsek model
Xinquan Chen ... Xianglin Bao
Journal of Intelligent & Fuzzy Systems | VOL. -
Xinquan Chen, et. al.Xinquan Chen ... Xianglin Bao
08 Sep 2023
Journal of Intelligent & Fuzzy Systems | VOL. -

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Big data: an optimized approach for cluster initialization

Abstract

Talk to us

Similar Papers

More From: Journal of Big Data