Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Comparison of Clustering Techniques in Text Documents in Portuguese

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Managing the vast amount of text data in the digital world is a complex challenge. An effective approach to tackle it is through the technique of text document clustering. This study evaluated the performance of three clustering algorithms — K-Means, Single Linkage, and Gaussian Mixture Model (GMM) — in clustering Brazilian Portuguese news articles using BERTimBau, a Portuguese variant of the BERT model, for preprocessing. Metrics such as accuracy, F1-score, Rand index, and Jaccard coefficient were used for evaluation. The results of these metrics indicated that Single Linkage achieved the best overall performance, surpassing K-Means and GMM in most of the evaluated criteria.

Similar Papers
  • Research Article
  • Cite Count Icon 6
  • 10.1080/03610918.2015.1082586
Effects of some design factors on the distribution of similarity indices in cluster analysis
  • Oct 23, 2015
  • Communications in Statistics - Simulation and Computation
  • Ahmed N Albatineh + 3 more

ABSTRACTThis article investigates the effects of number of clusters, cluster size, and correction for chance agreement on the distribution of two similarity indices, namely, Jaccard and Rand indices. Skewness and kurtosis are calculated for the two indices and their corrected forms then compared with those of the normal distribution. Three clustering algorithms are implemented: complete linkage, Ward, and K-means. Data were randomly generated from bivariate normal distributions with specified means and variance covariance matrices. Three-way ANOVA is performed to assess the significance of the design factors using skewness and kurtosis of the indices as responses. Test statistics for testing skewness and kurtosis and observed power are calculated. Simulation results showed that independent of the clustering algorithms or the similarity indices used, the interaction effect cluster size x number of clusters and the main effects of cluster size and number of clusters were found always significant for skewness and kurtosis. The three way interaction of cluster size x correction x number of clusters was significant for skewness of Rand and Jaccard indices using all clustering algorithms, but was not significant using Ward's method for both Rand and Jaccard indices, while significant for Jaccard only using complete linkage and K-means algorithms. The correction for chance agreement was significant for skewness and kurtosis using Rand and Jaccard indices when complete linkage method is used. Hence, such design factors must be taken into consideration when studying distribution of such indices.

  • Research Article
  • Cite Count Icon 6
  • 10.1007/s40815-016-0263-0
A Generalization of Rand and Jaccard Indices with Its Fuzzy Extension
  • Nov 14, 2016
  • International Journal of Fuzzy Systems
  • Chiou-Cherng Yeh + 1 more

The Jaccard and Rand indices are the best-known and used similarity measures. In general, the Jaccard index is relatively conservative, but the Rand index is relatively optimistic. In the paper, we make a generalization of Rand and Jaccard indices with its fuzzy extension. We first define a compromised weight to improve the Rand and Jaccard indices and provide the weight parameter selection. We then further advance this into fuzzy extension so that it can be used to measure similarities between fuzzy partitions and crisp reference partitions and those between fuzzy partitions and fuzzy reference partitions. Therefore, the proposed method is more flexible and reasonable to provide a useful way that can be applied in practical studies according to actual demands. Finally, we use simulation and make comparisons to complete more explanations and further discussions.

  • Research Article
  • Cite Count Icon 220
  • 10.1016/j.patrec.2006.11.010
A fuzzy extension of the Rand index and other related indexes for clustering and classification assessment
  • Jan 17, 2007
  • Pattern Recognition Letters
  • R.J.G.B Campello

A fuzzy extension of the Rand index and other related indexes for clustering and classification assessment

  • Research Article
  • Cite Count Icon 1
  • 10.3390/app16010540
Comparative Analysis of Clustering Algorithms for Unsupervised Segmentation of Dental Radiographs
  • Jan 5, 2026
  • Applied Sciences
  • Priscilla T Awosina + 2 more

In medical diagnostics and decision-making, particularly in dentistry where structural interpretation of radiographs plays a crucial role, accurate image segmentation is a fundamental step. One established approach to segmentation is the use of clustering techniques. This study evaluates the performance of five clustering algorithms, namely, K-Means, Fuzzy C-Means, DBSCAN, Gaussian Mixture Models (GMM), and Agglomerative Hierarchical Clustering for image segmentation. Our study uses two sets of real-world dental data comprising 140 adult tooth images and 70 children’s tooth images, including professionally annotated ground truth masks. Preprocessing involved grayscale conversion, normalization, and image downscaling to accommodate computational constraints for complex algorithms. The algorithms were accessed using a variety of metrics including Rand Index, Fowlkes-Mallows Index, Recall, Precision, F1-Score, and Jaccard Index. DBSCAN achieved the highest performance on adult data in terms of structural fidelity and cluster compactness, while Fuzzy C-Means excelled on the children dataset, capturing soft tissue boundaries more effectively. The results highlight distinct performance behaviours tied to morphological differences between adult and pediatric dental anatomy. This study offers practical insights for selecting clustering algorithms tailored to dental imaging challenges, advancing efforts in automated, label-free medical image analysis.

  • Research Article
  • Cite Count Icon 4
  • 10.1016/j.prevetmed.2015.03.002
Evaluation of a hierarchical ascendant clustering process implemented in a veterinary syndromic surveillance system
  • Mar 17, 2015
  • Preventive Veterinary Medicine
  • Isabelle Behaeghel + 8 more

Evaluation of a hierarchical ascendant clustering process implemented in a veterinary syndromic surveillance system

  • PDF Download Icon
  • Research Article
  • 10.64898/2026.01.02.26343331
ATN Classification and Machine-Learned Plasma Biomarker Phenotypes Reveal Distinct Alzheimer’s Pathology in a Population-Based Cohort
  • Feb 3, 2026
  • medRxiv
  • Emmanuel Fle Chea

Background:The ATN (Amyloid/Tau/Neurodegeneration) framework provides a theory-driven approach to Alzheimer’s disease (AD) classification using binary biomarker cutoffs, while unsupervised machine learning offers data-driven phenotyping. The concordance between these approaches in population-representative samples remains incompletely characterized.Objective:To compare plasma ATN classification with data-driven clustering methods and evaluate their associations with cognitive outcomes in a nationally representative cohort.Methods:We analyzed plasma biomarkers (Aβ42/40 ratio, p-tau181, NfL, GFAP) from 4,465 participants aged ≥51 years in the Health and Retirement Study 2016 Venous Blood Study. ATN profiles were classified using literature-based cutoffs. We applied k-means clustering, Gaussian mixture modeling, and variational autoencoder (VAE) dimensionality reduction to identify data-driven biomarker-defined subgroups. Agreement between ATN and clustering was quantified using adjusted Rand index (ARI) and normalized mutual information (NMI). Longitudinal analyses examined associations with cognitive decline over 4 years (2016–2020).Results:The analytic sample included 4,465 individuals (mean age 69.7±10.4 years; 58.7% female; 75.8% non-Hispanic White). ATN classification yielded 14 profiles, with A+/T−/N− (27.4%) and A−/T−/N− (22.6%) most prevalent (Figure 2). K-means clustering identified 4 optimal clusters with distinct biomarker signatures. Agreement between ATN and clusters was modest (ARI=0.119, NMI=0.113). Sensitivity analysis excluding GFAP from clustering reduced agreement substantially (ARI=0.03 vs 0.119 with GFAP, −74.5% decrease), demonstrating that GFAP accounts for most of the observed concordance between clustering and ATN classification, with only one-third arising from the shared three biomarkers.[Table S12] Additional sensitivity analyses confirmed that k=4 provides finer biomarker resolution than k=3 by retaining biomarker extreme subgroups[Table S13], and that Cluster 4 represents a stable biological structure across distance metrics[Table S14] despite its small size. Cluster 1 (n=51, 1.2%) showed severe pathology; Cluster 3 (n=3,479, 78.6%) represented the largest and most heterogeneous group, encompassing the broad spectrum of minimal to moderate pathology across all ATN profiles; Cluster 4 (n=14, 0.3%) represented a small but stable non AD biomarker defined subgroup (Jaccard=0.779). The VAE revealed a localized nonlinear structure. Silhouette values in the latent space are not directly comparable to clustering silhouettes, but the VAE embedding showed clearer local separation, whereas PCA explained more variance (67.1%). Both ATN and clusters predicted 4-year cognitive decline (ATN R2=0.024, p<0.001; Clusters R2=0.019, p<0.001).Conclusions:Theory driven ATN classification and data driven biomarker phenotyping capture partially overlapping but largely distinct information. Modest concordance (ARI=0.119) reflects GFAP’s contribution to shared structure, with most alignment arising from GFAP rather than from the three ATN biomarkers alone (ARI=0.03). The primary source of discordance remains the binary versus continuous representation of biomarker variation. Sensitivity analyses showed that k=4 provides finer biomarker resolution than k=3, and that Cluster 4 represents a small but reproducible biomarker defined subgroup. Both approaches predicted cognitive decline with modest effect sizes (R2=1.9–2.4%), consistent with population based studies. Integrating theory driven and data driven frameworks may support a more comprehensive characterization of AD related pathology in population research.

  • Research Article
  • Cite Count Icon 2
  • 10.5121/mlaij.2024.11402
Comparative Analysis of Clustering Algorithms on Synthetic Circular Patters Data
  • Dec 28, 2024
  • Machine Learning and Applications: An International Journal
  • Hardev Ranglani

Clustering algorithms play a pivotal role in discovering hidden patterns in unlabeled data, but their performance varies significantly across datasets with complex geometries. This paper explores the performance of various clustering techniques in identifying distinct circular clusters within the Synthetic Circle Data Set, a benchmark dataset designed to test algorithms on non-linear structures. We evaluate popular clustering methods, including k-means, DBSCAN, Gaussian Mixture Models, hierarchical clustering, and emerging techniques like Self Organizing Maps, Mean Shift Clustering and Spectral Clustering. Using metrics such as Adjusted Rand Index (ARI), Normalized Mutual Information (NMI), and Silhouette Score, along with detailed visualizations, we systematically compare the algorithms’ ability to recover the true circle-based clusters without prior labels. Our findings highlight the strengths and limitations of each method, revealing that density- and graph-based algorithms consistently outperform traditional techniques like k-means in handling circular patterns.

  • Research Article
  • 10.1063/5.0214782
Ensemble of fuzzy c-means clustering algorithm based on cluster label uniformity for MRI brain image segmentation
  • Jul 1, 2025
  • AIP Advances
  • Anup Kumar Mallick + 3 more

This paper proposes a novel ensemble method with multiple runs of the fuzzy c-means (FCM) clustering algorithm as the base learners. However, the major challenge in the ensemble of different runs of FCM lies in the labeling of clusters in each run. As there is no uniformity for numbering the cluster labels, the data points belonging to close clusters may be assigned to different clusters across multiple runs of FCM. Hence, assembling the cluster solutions of different runs of FCM becomes difficult. This challenge has been addressed in this study by proposing a novel method to ensemble multiple runs of FCM. In the proposed method, the cluster labels of different runs are renumbered based on proximity among the clusters, making it easy to ensemble cluster solutions in multiple runs of FCM. The class labels of the non-ensemble data points are determined by using three state-of-the-art classifiers, namely, k-nearest neighbors, support vector machine, and artificial neural network. The proposed method is applied for segmenting magnetic resonance imaging (MRI) of the human brain. The division of MRI brain images into distinct tissue classes is a crucial aspect of neurological disease and clinical research. The performance of the proposed method in segmenting the MRI brain images is compared to the Gaussian mixture models based on the values of three performance measures, namely, the Rand index, adjusted Rand index, and Minkowski score. The simulation results demonstrate the supremacy of the proposed method over the Gaussian mixture models.

  • Research Article
  • Cite Count Icon 1
  • 10.5351/kjas.2007.20.1.167
고차원 (유전자 발현) 자료에 대한 군집 타당성분석 기법의 성능 비교
  • Mar 31, 2007
  • Korean Journal of Applied Statistics
  • Yun-Kyoung Jeong + 1 more

유전자 발현 자료(gene expression data)는 전형적인 고차원 자료이며, 이를 분석하기 위한 여러 가지 군집 알고리즘(clustering algorithm)과 군집 결과들을 검증하는 군집타당성분석 기법(cluster validation technique)이 제안되고 있지만, 이들 군집 타당성을 분석하는 기법의 성능에 대한 비교, 평가는 매우 드물다. 본 논문에서는 저차원의 모의실험 자료와 실제 유전자 발현 자료에 대하여 군집 타당성분석 기법들의 성능을 비교하였으며, 그 결과 내적 측도에서는 Dunn 지수, Silhouette 지수 순으로 뛰어났고 외적 측도에서는 Jaccard 지수가 성능이 가장 우수한 것으로 평가되었다. Many clustering algorithms and cluster validation techniques for high-dimensional gene expression data have been suggested. The evaluations of these cluster validation techniques have, however, seldom been implemented. In this paper we compared various cluster validity indices for low-dimensional simulation data and real gene expression data, and found that Dunn's index is the most effective and robust, Silhouette index is next and Davies-Bouldin index is the bottom among the internal measures. Jaccard index is much more effective than Goodman-Kruskal index and adjusted Rand index among the external measures.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 16
  • 10.46481/jnsps.2023.1364
An Empirical Study on Anomaly Detection Using Density-based and Representative-based Clustering Algorithms
  • Apr 19, 2023
  • Journal of the Nigerian Society of Physical Sciences
  • Gerard Shu Fuhnwi + 3 more

In data mining, and statistics, anomaly detection is the process of finding data patterns (outcomes, values, or observations) that deviate from the rest of the other observations or outcomes. Anomaly detection is heavily used in solving real-world problems in many application domains, like medicine, finance , cybersecurity, banking, networking, transportation, and military surveillance for enemy activities, but not limited to only these fields. In this paper, we present an empirical study on unsupervised anomaly detection techniques such as Density-Based Spatial Clustering of Applications with Noise (DBSCAN), (DBSCAN++) (with uniform initialization, k-center initialization, uniform with approximate neighbor initialization, and $k$-center with approximate neighbor initialization), and $k$-means$--$ algorithms on six benchmark imbalanced data sets. Findings from our in-depth empirical study show that k-means-- is more robust than DBSCAN, and DBSCAN++, in terms of the different evaluation measures (F1-score, False alarm rate, Adjusted rand index, and Jaccard coefficient), and running time. We also observe that DBSCAN performs very well on data sets with fewer number of data points. Moreover, the results indicate that the choice of clustering algorithm can significantly impact the performance of anomaly detection and that the performance of different algorithms varies depending on the characteristics of the data. Overall, this study provides insights into the strengths and limitations of different clustering algorithms for anomaly detection and can help guide the selection of appropriate algorithms for specific applications.

  • Conference Article
  • Cite Count Icon 7
  • 10.1109/cinc.2009.214
Comparison of Cluster Ensembles Methods Based on Hierarchical Clustering
  • Jun 1, 2009
  • Kai Li + 2 more

Cluster ensembles method is considered as a robust and accurate alternative to single clustering runs. It mainly consists of both generation of individual member and fusion methods. In this paper, we study the cluster ensembles where individual members are obtained based on k-means clustering algorithm and fusion method of hierarchical clustering is used. Three consensus functions, which are single linkage, complete linkage and average linkage, respectively, is studied and discussed in hierarchical clustering fusion. For evaluating performance of cluster ensembles, adjusted rand index is considered. Experimental results show that performance of cluster ensembles with the average linkage is superior to one with single linkage and complete linkage. Moreover, we also study the relationship between accuracy and ensemble size of the three methods.

  • Research Article
  • Cite Count Icon 6
  • 10.3390/electronics14071429
Lasso-Based k-Means++ Clustering
  • Apr 1, 2025
  • Electronics
  • Shazia Parveen + 1 more

Clustering is a powerful and efficient technique for pattern recognition which improves classification accuracy. In machine learning, it is a useful unsupervised learning approach due to its simplicity and efficiency for clustering applications. The curse of dimensionality poses a significant challenge as the volume of data increases with rapid technological advancement. It makes traditional methods of analysis inefficient. Sparse clustering is essential for efficiently processing and analyzing large-scale, high-dimensional data. They are designed to handle and process sparse data efficiently since most elements are zero or lack information. In data science and engineering applications, they play a vital role in taking advantage of the natural sparsity in data to save computational resources and time. Motivated by recent sparse k-means and k-means++ algorithms, we propose two novel Lasso-based k-means++ (Lasso-KM++) clustering algorithms, Lasso-KM1++ and Lasso-KM2++, which incorporate Lasso regularization to enhance feature selection and clustering accuracy. Both Lasso-KM++ algorithms can shrink the irrelevant features towards zero, and select relevant features effectively by exploring better clustering structures for datasets. We use numerous synthetic and real datasets to compare the proposed Lasso-KM++ with k-means, k-means++ and sparse k-means algorithms based on the six performance measures of accuracy rate, Rand index, normalized mutual information, Jaccard index, Fowlkes–Mallows index, and running time. The results and comparisons show that the proposed Lasso-KM++ clustering algorithms actually improve both the speed and the accuracy. They demonstrate that our proposed Lasso-KM++ algorithms, especially for Lasso-KM2++, outperform existing methods in terms of efficiency and clustering accuracy.

  • Research Article
  • Cite Count Icon 19
  • 10.1007/s41870-019-00406-7
Partitioning and hierarchical based clustering: a comparative empirical assessment on internal and external indices, accuracy, and time
  • Nov 29, 2019
  • International Journal of Information Technology
  • Syed Imtiyaz Hassan + 3 more

Clustering is an unsupervised data mining technique where exploration is done with little knowledge of data classes. Its aim is to recognize the hidden information from the data for effective decision-making. Though many clustering algorithms has already been implemented till date, still it is an active topic of research for data mining. Researcher’s attempts to explore, compare, evaluate, and improve the different clustering algorithms available, for specialized situation and context. The purpose of all these efforts are to refine and propose improved version of algorithm after statistical evaluation by different metrices. The present research is an attempt to analysis empirically, the partitioning based clustering algorithms and hierarchical based clustering algorithm; by conducting extensive experiments. Both algorithms effectiveness has been measured through external and internal validity indices and Pearson’s correlation distance function using anatomized experiments. The parameters of evaluation that have been taken into consideration; for Internal Indices: Silhouette Index, Davies-Bouldin Validity Index and Calinski-Harabasz index; for external indices: Jaccard index, Rand Index, Entropy and Normalized Mutual Information. The other parameters of evaluation are accuracy and time of execution. Based on the experiments it may be concluded that K-means algorithm produces more promising result than hierarchical algorithm except in accuracy.

  • Research Article
  • Cite Count Icon 9
  • 10.1109/tim.2023.3279913
Soft Sensing of NOx Emissions From Thermal Power Units Based on Adaptive GMM Two-Step Clustering Algorithm and Ensemble Learning
  • Jan 1, 2023
  • IEEE Transactions on Instrumentation and Measurement
  • Ze Dong + 3 more

The accurate measurement of NOx concentration is the basis of accurate ammonia injection in selective catalytic reduction (SCR) system of thermal power unit. Excessive or too little ammonia injection will cause ammonia escape or exceeding the standard of NOx emission, and cause environmental pollution. There exist a large measurement lag and errors during the purging process in the measurement of NOx concentration in the SCR system of thermal power units. Therefore, to enable the intelligent control of the SCR system, it is necessary to carry out soft sensing of NOx emissions. This paper proposes a soft sensing algorithm for NOx emissions of thermal power units based on an adaptive Gaussian Mixture Model (GMM) two-step clustering algorithm and ensemble learning. Adaptive GMM two-step clustering (AGTSC) algorithm is used to softly divide the historical data into working conditions. Temporal Pattern Attention Long Short-Term Memory (TPA-LSTM) is also adopted to construct individual learners for each working condition. Here we use the parameter regression algorithm to train the combiner based on the output of the individual learner and the membership signal of the working condition as the input of the combiner. The clustering algorithm is tested with three datasets. The results show that this algorithm addresses the shortcomings of GMM including not being applicable to the clustering of non-convex datasets and manual selection of the number of clusters. The soft sensing algorithm is verified by using the historical data of the SCR system of 1000 <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">MW</i> thermal power units. The results show that the proposed method has higher accuracy and stronger generalization ability than the conventional method, hence providing an effective method for soft sensing of NOx emissions of thermal power units.

  • Research Article
  • Cite Count Icon 6
  • 10.1016/j.camwa.2007.07.019
Applying the extended mass-constraint EM algorithm to image retrieval
  • Apr 2, 2008
  • Computers &amp; Mathematics with Applications
  • Daan He + 2 more

Applying the extended mass-constraint EM algorithm to image retrieval

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant