A Multivariate Fuzzy Weighted K-Modes Algorithm with Probabilistic Distance for Categorical Data

Ren-Jieh Kuo,Thi Phuong Quyen Nguyen,Maya Cendana,Ferani E Zulvia

doi:10.5614/itbj.ict.res.appl.2023.18.2.2

Ren-Jieh Kuo, Thi Phuong Quyen Nguyen + Show 2 more

Open Access

https://doi.org/10.5614/itbj.ict.res.appl.2023.18.2.2

Copy DOI

Export

Save

Cite

Journal: Journal of ICT Research and Applications	Publication Date: Sep 30, 2024
License type: cc-by-nd

Abstract
Full-Text
Similar Papers

Abstract

Listen

Data clustering is a data mining approach that assigns similar data to the same group. Traditionally, cluster similarity considers all attributes equally, but in real-world applications, some attributes may be more important than others. Therefore, this study proposes an algorithm that utilizes multivariate fuzzy weighting to demonstrate the varying importance of each attribute, using a Gini impurity measure for weight assignment. Additionally, the proposed algorithm implements probabilistic distance to reduce sensitivity to noise. Probabilistic distance offers more detailed information and better interpretation than Hamming distance, which ignores attribute positions. Probabilistic distance utilizes information about the attribute’s position within and between clusters. This enhances clustering performance by creating clusters with more similar attributes. Therefore, the proposed Multivariate Fuzzy Weighted K-Modes with Probabilistic Distance for Categorical Data (MFWKM-PD) algorithm, based on the multivariate fuzzy K-modes algorithm, not only considers detailed membership calculations but also considers the varying contributions of attributes and their positions in distance calculation. This study evaluated the proposed MFWKM-PD using several benchmark datasets. The experiments validated that the proposed MFWKM-PD shows promising results compared to other algorithms in terms of accuracy, NMI, and ARI.

Full Text

Published Version

View

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

A Multivariate Fuzzy Weighted K-Modes Algorithm with Probabilistic Distance for Categorical Data

Abstract

Published Version

Talk to us

Similar Papers

More From: Journal of ICT Research and Applications

Lead the way for us

Similar Papers

A Novel Consensus Fuzzy K-Modes Clustering Using Coupling DNA-Chain-Hypergraph P System for Categorical Data
Zhenni Jiang ... Xiyu Liu
Processes | VOL. 8
Zhenni Jiang, et. al.Zhenni Jiang ... Xiyu Liu
21 Oct 2020
Processes | VOL. 8

Application of metaheuristic based fuzzy K-modes algorithm to supplier clustering
R.J Kuo ... Yuliana Potti
Computers & Industrial Engineering | VOL. 120
R.J Kuo, et. al.R.J Kuo ... Yuliana Potti
28 Apr 2018
Computers & Industrial Engineering | VOL. 120

Categorical Data Clustering Method Based on Improved Fruit Fly Optimization Algorithm
Dong Li ... Yan Zhang
-
Dong Li, et. al.Dong Li ... Yan Zhang
01 Jan 2019
01 Jan 2019

Genetic intuitionistic weighted fuzzy k-modes algorithm for categorical data
R.J Kuo ... Thi Phuong Quyen Nguyen
Neurocomputing | VOL. 330
R.J Kuo, et. al.R.J Kuo ... Thi Phuong Quyen Nguyen
15 Nov 2018
Neurocomputing | VOL. 330

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

A Multivariate Fuzzy Weighted K-Modes Algorithm with Probabilistic Distance for Categorical Data

Abstract

Published Version

Talk to us

Similar Papers

More From: Journal of ICT Research and Applications