Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Breast cancer diagnosis based on feature extraction using a hybrid of K-means and support vector machine algorithms

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Breast cancer diagnosis based on feature extraction using a hybrid of K-means and support vector machine algorithms

Similar Papers
  • Research Article
  • 10.32520/stmsi.v13i4.4113
Breast Cancer Classification based on Ultrasound Images using the Support Vector Machine (SVM) Algorithm
  • Jul 29, 2024
  • SISTEMASI
  • Nurazmi Aprilia + 1 more

According to statistics from the Global Burden of Cancer Study (Globocon) of the World Health Organization (WHO), cancer, particularly breast cancer, is a severe health issue in Indonesia with 68,858 new cases and 22,000 deaths recorded in 2020. Ultrasonography (USG) technology is acknowledged as one of the potentials to support early detection, which is vital in reducing mortality from breast cancer. This study focuses on classifying ultrasound images using the Support Vector Machine (SVM) algorithm, GLCM feature extraction, Min-Max normalization, and Mutual Information with SelectKBest Feature Selection. From several experiments using the SVM algorithm with various combinations of parameter values that have been set and different Tests, namely using a Train/Test Split with a proportion of 80/20 and K-Fold Cross Validation, it shows that the SVM algorithm is capable of classifying ultrasound images of breast cancer. into two categories (Benign Tumor and Malignant Tumor) with the same maximum accuracy of 79% after applying the SMOTE Balancing Data technique or without using the Balancing Data technique. As a result, the Support Vector Machine (SVM) algorithm has the potential to be an effective model for identifying breast cancer ultrasound images, both on data from the original set that has not been balanced and data from the set that has been balanced.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 2
  • 10.1155/2021/5289128
Value of Magnetic Resonance Imaging Features in Diagnosis and Treatment of Breast Cancer under Intelligent Algorithms
  • Nov 3, 2021
  • Scientific Programming
  • Shuang Liu + 4 more

This study was to analyze the clinical application value of magnetic resonance imaging (MRI) image features based on intelligent algorithms in the diagnosis and treatment of breast cancer and to provide an effective reference assessment for breast cancer diagnosis. The MRI diagnosis model (ACO-MRI) based on the ant colony algorithm (ACO) was proposed, which was compared with the diagnosis methods based on support vector machine (SVM) and proximity (KNN) algorithm, and the proposed algorithm was applied to MRI images to diagnose breast cancer. The results showed that the accuracy, sensitivity, and specificity of the ACO-MRI model were greater than those of the KNN and SVM algorithm. Moreover, the specificity was statistically considerable compared with the two algorithms of KNN and SVM ( P < 0.05 ). By comparing 1/5 number of ants and the average gray path of the ACO-MRI model under 1/8 number of ants, it was found that the average gray path value of 1/8 number of ants was greatly higher than the average gray path value of 1/5 number of ants ( P < 0.05 ). The differences in the overall distribution of breast MRI imaging features among Luminal A, Luminal B, HER-2 overexpression, and TN were compared. There were considerable differences in the overall distribution of the three breast MRI imaging features of the boundaries, morphology, and enhancement methods among the four groups ( P < 0.05 ). In short, MRI image based on the intelligent algorithm ACO-MRI diagnosis model can effectively improve the diagnosis effect of breast cancer. Its image feature boundaries, morphology, and enhancement methods had good imaging features in the diagnosis of breast cancer.

  • Conference Article
  • Cite Count Icon 18
  • 10.1109/mcsoc51149.2021.00057
The Role of Linear Discriminant Analysis for Accurate Prediction of Breast Cancer
  • Dec 1, 2021
  • Egwom Onyinyechi Jessica + 3 more

With the recent advances in clinical technologies, a huge amount of data has been accumulated for breast cancer diagnosis. Extracting information from the data to support the clinical diagnosis of breast cancer is a tedious and time-consuming task. The use of machine learning and data mining techniques has significantly changed the whole process of a breast cancer diagnosis. In this research, a prediction model for breast cancer prediction has been developed using features extracted from individual medical screening and tests. To overcome the problem of overfitting and obtain a good prediction accuracy, a Linear Discriminant Analysis (LDA) is applied for the extraction of useful features. This is done to reduce the number of features in the experimental dataset. The proposed model can create new features from the existing features and then get rid of the original features. The newly created features were able to summarize the necessary information contained initially in the original set of features. LDA was chosen because of its usefulness in detecting whether a set of features is worthwhile in predicting breast cancer. In addition to LDA, the proposed model uses Support Vector Machine (SVM) for accurate prediction, hence the name LDA-SVM prediction model. Based on 5-fold cross-validation, the proposed model yields an accuracy of 99.2%, precision of 98.0%, and Recall of 99.0% when tested on the Wisconsin Diagnostic Breast Cancer (WDBC) dataset from the University of California- Irvine machine learning repository. Therefore, SVM shows high efficiency in handling classification problems when combined with feature extraction techniques.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 1
  • 10.1155/2023/4097660
Study on Gene Splicing Site Recognition Based on Particle Swarm Optimization Twin Support Vector Machine Algorithm for Smart Healthcare
  • Apr 21, 2023
  • Wireless Communications and Mobile Computing
  • Fuquan Zhang + 6 more

Gene splicing site recognition is a very important research topic in smart healthcare. Gene splicing site recognition is of great significance, not only for the large-scale and high-quality computational annotation of genomes but also for the analysis and recognition of the gene sequences evolutionary process. It is urgent to study a reliable and effective algorithm for gene splice site recognition. Traditional Twin Support Vector Machine (TWSVM) algorithm has advantages in solving small-sample, nonlinear, and high-dimensional problems, but it cannot deal with parameter selection well. To avoid the blindness of parameter selection, particle swarm optimization algorithm was used to find the optimal parameters of twin support vector machine. Therefore, a Particle Swarm Optimization Twin Support Vector Machine (PSO-TWSVM) algorithm for gene splicing site recognition was proposed in this paper. The proposed algorithm was compared with traditional Support Vector Machine algorithm, TWSVM algorithm, and Least Squares Support Vector Machine algorithm. The comparison results show that the positive sample recognition rate, negative sample recognition rate, and correlation coefficient (CC) of the proposed algorithm are the best among the four different support vector machine algorithms. The proposed algorithm effectively improves the recognition rate and the accuracy of splice sites. The comparison experiments verify the feasibility of the proposed algorithm.

  • Conference Article
  • Cite Count Icon 6
  • 10.1063/5.0110987
Comparison of tuberculosis disease classification using support vector machine and Naive Bayes algorithm
  • Jan 1, 2023
  • AIP conference proceedings
  • Yogiek Indra Kurniawan + 4 more

Tuberculosis (TB) is one of the biggest causes of death in the world. Some several symptoms and factors can be used to identify TB disease. One way that can be used is the classification method in data mining. This study compared the classification of TB disease using the Support Vector Machine (SVM) and Naive Bayes Algorithm. The research started by collecting data, then divided them into 13 independent variables and a dependent variable. After that, SVM and Naïve Bayes are implemented to classify the data. Based on the test results for each algorithm, with giving the amount of training data greater can give impact to the higher precision, recall, and accuracy values of the two algorithms used. In addition, for the case of Tuberculosis disease classification, the value of precision, recall, and accuracy in the Support Vector Machine algorithm is always higher than the Naive Bayes algorithm in all provided data.

  • Research Article
  • 10.1088/1742-6596/1539/1/012007
Determination of Granting Appropriateness Credit at “Daruzzakah Rensing” Cooperative Using the Support Vector Machine (SVM) Algorithm
  • May 1, 2020
  • Journal of Physics: Conference Series
  • Yahya + 1 more

Credit is the main product of savings and loan cooperatives to increase profitability. The greater the credit issued, the greater the benefits obtained by cooperatives. Each cooperative will package credit products in such a way as to attract the attention of every customer. However, cooperatives can find problems in the process of lending, such as the “Daruzzakah Rensing” Cooperative located in “Desa Rensing, Kecamatan Sakra Barat, Lombok Timur-NTB-Indonesia”. The main products of the Cooperative “Daruzzakah Rensing” are savings and loans. In distributing credit, the cooperative always decides based on statistical data. This data is sometimes not useful if the supporting methods used to predict and classify the data are not appropriate. Therefore, this research requires a method that can classify and predict problematic and non-problematic customers. To answer this question, using the SVM (Support Vector Machine) algorithm to find out the level of accuracy in analyzing creditworthiness proposed by prospective debtors. The SVM algorithm is used to predict, classify, evaluate, and analyze credit. From the results of data processing carried out using the SVM algorithm (Support Vector Machine), it can be categorized as an excellent method, with an accuracy of 90.42% and AUC at 0.957. Accuracy of 90.42% means the SVM algorithm can provide decisions about feasible or not feasible in granting credit to customers who apply for loans.

  • Research Article
  • Cite Count Icon 72
  • 10.1007/s13042-017-0741-1
Feature extraction by PCA and diagnosis of breast tumors using SVM with DE-based parameter tuning
  • Nov 9, 2017
  • International Journal of Machine Learning and Cybernetics
  • Luanyi Yang + 1 more

Breast cancer is the second most common cause of death among the women worldwide, whereas the early detection may well lead to a longer survival or even full recovery. With the development of clinical technologies, massive tumor feature data become available to be collected and meanwhile many machine learning techniques have been introduced to support doctors in diagnostic decision-making process. In this paper, we develop a support vector machine (SVM) based diagnosing system which mainly consists of three stages. For the first stage, principal component analysis is implemented to eliminate the redundant information and extract representive patterns out of the original data. This procedure reduces the feature space dimension which cuts down the computational complexity significantly. In the second stage, we search the optimal parameter values for SVM using the differential evolution algorithm. At last, a classifier is trained to differentiate the incoming tumors. In order to objectively and comprehensively evaluate the classifier’s performance, a series of indices are considered simultaneously such as classification accuracy, sensitivity, specificity and area under receiver operating characteristic curves. In comparison with K-nearest neighbor, random forest, bagging, naive bayes, decision tree and other classificaiton approaches, the proposed method presents a superior performance when tested on the Wisconsin Diagnostic Breast Cancer (WDBC) data set from the University of California with fivefold cross validation.

  • Research Article
  • 10.54254/2755-2721/5/20230555
Comparison between Bayesian and SVM model for breast cancer risk prediction
  • May 31, 2023
  • Applied and Computational Engineering
  • Jiatong Jiang + 3 more

Breast cancer ranks first in the incidence of cancer worldwide. It is a kind of malignant disease with incidence increasing year by year. Recently, machine learning algorithms widely used in cancer prediction and other fields to relieve the burden of doctors and accelerate the diagnostic process. In this work, two representative models, the support vector machine (SVM) and Bayesian classification algorithm are leveraged for breast cancer risk prediction. It is found that the effect and accuracy of the two algorithms are very different when they are used for breast cancer prediction. However, there is still a research gap in the comparison of the two algorithms in practical application. Therefore, the research topic of this paper is the comparison between the Bayesian classification algorithm and the SVM algorithm in breast cancer prediction. The research methods of this paper are as follows: firstly, the dataset is collected, then applying the Bayesian classification to process the dataset, and then the SVM algorithm is used to process the dataset. Finally, the processing results of the two algorithms are compared to comprehensively analyze and compare which algorithm is more efficient and more suitable for actual breast cancer prediction. The test results support that the SVM outperforms the Bayesian classification algorithm in the actual target tracking problem. Therefore, it is suggested to choose the SVM classification algorithm in the actual target tracking problem to boost the accuracy and efficiency of prediction results to the greatest extent.

  • Research Article
  • 10.18137/cardiometry.2022.25.878884
Analysis and Comparison for Innovative Prediction Technique of Breast Cancer Tumor using k Nearest Neighbor Algorithm over Support Vector Machine Algorithm with Improved Accuracy
  • Feb 14, 2023
  • CARDIOMETRY
  • Ch Srinivasulureddy + 1 more

Aim: The main objective of this study is to compare the efficiency of the k-Nearest Neighbor (KNN) and Support vector machine (SVM) algorithms in detecting breast cancer tumors and to examine their improved accuracy, sensitivity, and precision. Materials and Methods: The data for the research of Innovative breast cancer prediction using machine learning algorithms is taken from UCI Machine Learning Repository. The sample size of the innovative technique involves two groups KNN (N=20) and SVM (N=20) according to clincalc.com by keeping alpha error-threshold at 0.05, confidence interval at 95%, enrollment ratio as 0:1, and power at 80%. The accuracy, sensitivity, and precision are calculated using MATLAB software. Result: Accuracy (%), sensitivity (%), precision (%) are compared using SPSS software using an independent sample t-test tool. The accuracy of the k-Nearest Neighbor is 93.38% (p<0.001) while the accuracy of the Support vector machine is 97.50%. The sensitivity rate is 90.85% (p<0.001) for k-Nearest Neighbor whereas the results of Support vector machine sensitivity is 95.83%. The precision of k-Nearest Neighbor is 98.48% (p<0.001) whereas the results of Support vector machine precision is 100%. Conclusion: The support vector machine algorithm appears to have performed better than the k-Nearest Neighbor with improved accuracy in Innovative breast cancer prediction.

  • Research Article
  • 10.55606/jupumi.v5i1.4347
Klasifikasi Jenis Kelamin Berbasis Citra Mata Menggunakan Algoritma Support Vector Machine
  • Jan 15, 2026
  • Jurnal Publikasi Manajemen Informatika
  • Abdurrahman Afifi + 1 more

This study aims to develop a gender classification model based on eye images using the Support Vector Machine (SVM) algorithm. The dataset consists of 13,499 eye images divided into two classes: male and female. The methodology includes preprocessing by converting images to grayscale and resizing them to 64×64 pixels, followed by feature extraction using raw pixel representation resulting in a 4,096-dimensional vector. The data is split into 80% for training and 20% for testing, and SVM parameters are optimized using grid search with 5-fold cross-validation. The SVM model employs an RBF kernel with parameters C=10 and gamma='scale'.Evaluation is carried out using accuracy, precision, recall, F1-score metrics, and a confusion matrix. A decision boundary is visualized using PCA to analyze data separability. The results show excellent performance with 99.96% accuracy, 100.00% precision, 99.95% recall, and 99.98% F1-score. The confusion matrix indicates near-perfect classification, with 648 male samples and 2,051 female samples correctly classified without misprediction. This study demonstrates that the SVM algorithm, even with simple preprocessing, can achieve high accuracy in gender classification based on eye images, showing strong potential for practical implementation in biometric systems

  • Research Article
  • Cite Count Icon 1
  • 10.46336/ijqrm.v6i2.1011
Comparison of Random Forest and SVM Algorithms in Classification of Diabetic Retinopathy Based on Fundus Image Texture Features
  • Jun 10, 2025
  • International Journal of Quantitative Research and Modeling
  • Renda Sandi Saputra + 1 more

Diabetic Retinopathy (DR) is a microangiopathic complication of diabetes mellitus that can cause visual impairment to permanent blindness. Early detection of DR is essential to prevent disease progression, but conventional methods require time, cost, and expertise that are not always available. This study aims to compare the performance of the Random Forest (RF) and Support Vector Machine (SVM) algorithms in DR classification based on texture features extracted from retinal fundus images. The dataset used consists of 3,000 retinal fundus images obtained from the Kaggle platform, divided into 2,400 training data and 600 test data. Image preprocessing includes conversion to grayscale, resizing to a resolution of 128×128 pixels, and normalization. Feature extraction is performed using a combination of Local Binary Pattern (LBP) and Gray Level Co-occurrence Matrix (GLCM) to produce a 14-dimensional feature vector. Performance evaluation uses accuracy, precision, recall, F1-score, ROC curve, and 5-fold cross-validation metrics. The results showed that Random Forest significantly outperformed SVM with an accuracy of 96% compared to 64%, an AUC value of 0.99 compared to 0.72, and an average cross-validation accuracy of 94.5% compared to 63.42%. Random Forest also showed balanced performance in both classes with precision, recall, and F1-score of 0.96, while SVM experienced classification imbalance especially in the disease class. This study proves that Random Forest is a more optimal algorithm for an automatic DR detection system based on fundus image texture features and can support increasing the accessibility of DR screening in areas with limited specialist medical personnel.

  • Research Article
  • 10.22146/jnteti.v12i4.5125
Indonesian Society’s Sentiment Analysis Against the COVID-19 Booster Vaccine
  • Nov 22, 2023
  • Jurnal Nasional Teknik Elektro dan Teknologi Informasi
  • Dionisia Bhisetya Rarasati + 2 more

The COVID-19 pandemic is still occurring in various countries, including Indonesia. This pandemic is caused by the coronavirus, which has mutated into multiple virus variants, such as Delta and Omicron. As of 9 February 2022, 4,626,936 people were confirmed positive for COVID-19 in Indonesia. This number continues to rise. The Indonesian government has prevented the spread of these virus variants by introducing booster vaccines to the public. However, this vaccination program has caused various sentiments among Indonesians. To optimize efforts to combat COVID-19, the government needs to know these sentiments immediately. Based on these problems, the researcher proposes the application of machine learning technology to develop a system that can analyze the sentiments of the Indonesians toward the booster vaccine. This research has several stages: data collection, data labeling, text preprocessing, feature extraction, and application of the support vector machine (SVM) algorithm using various kernels, namely the linear kernel, Gaussian radial basis function (RBF) kernel, and polynomial kernel. Furthermore, the results of the system were tested for accuracy using a 10-fold cross validation and confusion matrix. The dataset used was 681 tweets with the hashtag “vaksinbooster.” The dataset consists of two classes: negative (0) and positive (1). The results showed that the data were positive for the booster vaccine, as evidenced by the higher number of positive tweets, with 554 data, compared to 127 negative tweets. In addition, the dataset was divided into training data of 545 and testing data of 136. In addition, the test results of this study revealed that the SVM algorithm with the polynomial kernel, which was evaluated with 10-fold cross validation, yielded the highest level of accuracy, namely 79.22%.

  • Research Article
  • Cite Count Icon 2
  • 10.1088/1757-899x/319/1/012062
Gradient Evolution-based Support Vector Machine Algorithm for Classification
  • Mar 1, 2018
  • IOP Conference Series: Materials Science and Engineering
  • Ferani E Zulvia + 1 more

This paper proposes a classification algorithm based on a support vector machine (SVM) and gradient evolution (GE) algorithms. SVM algorithm has been widely used in classification. However, its result is significantly influenced by the parameters. Therefore, this paper aims to propose an improvement of SVM algorithm which can find the best SVMs’ parameters automatically. The proposed algorithm employs a GE algorithm to automatically determine the SVMs’ parameters. The GE algorithm takes a role as a global optimizer in finding the best parameter which will be used by SVM algorithm. The proposed GE-SVM algorithm is verified using some benchmark datasets and compared with other metaheuristic-based SVM algorithms. The experimental results show that the proposed GE-SVM algorithm obtains better results than other algorithms tested in this paper.

  • Research Article
  • Cite Count Icon 22
  • 10.4015/s1016237221500204
BREAST CANCER DETECTION USING RSFS-BASED FEATURE SELECTION ALGORITHMS IN THERMAL IMAGES
  • Mar 9, 2021
  • Biomedical Engineering: Applications, Basis and Communications
  • Nazila Darabi + 2 more

Breast cancer is a common cancer in female. Accurate and early detection of breast cancer can play a vital role in treatment. This paper presents and evaluates a thermogram based Computer-Aided Detection (CAD) system for the detection of breast cancer. In this CAD system, the Random Subset Feature Selection (RSFS) algorithm and hybrid of minimum Redundancy Maximum Relevance (mRMR) algorithm and Genetic Algorithm (GA) with RSFS algorithm are utilized for feature selection. In addition, the Support Vector Machine (SVM) and k-Nearest Neighbors (kNN) algorithms are utilized as classifier algorithm. The proposed CAD system is verified using MATLAB 2017 and a dataset that is composed of breast images from 78 patients. The implementation results demonstrate that using RSFS algorithm for feature selection and kNN and SVM algorithms as classifier have accuracy of 85.36% and 75%, and sensitivity of 94.11% and 79.31%, respectively. In addition, using hybrid GA and RSFS algorithm for feature selection and kNN and SVM algorithms as classifier have accuracy of 83.87% and 69.56%, and sensitivity of 96% and 81.81%, respectively, and using hybrid mRMR and RSFS algorithms for feature selection and kNN and SVM algorithms as classifier have accuracy of 77.41% and 73.07%, and sensitivity of 98% and 72.72%, respectively.

  • Research Article
  • Cite Count Icon 41
  • 10.1142/s219688882150007x
Breast Cancer Detection Based on Feature Selection Using Enhanced Grey Wolf Optimizer and Support Vector Machine Algorithms
  • Nov 5, 2020
  • Vietnam Journal of Computer Science
  • Sunil Kumar + 1 more

Breast cancer is the leading cause of high fatality among women population. Identification of the benign and malignant tumor at correct time plays a critical role in the diagnosis of breast cancer. In this paper, an attempt has been made to extract the valuable information by selecting the relevant features using our proposed EGWO-SVM (enhanced grey wolf optimization-support vector machine) approach. Grey wolf optimizer (GWO) has gained a lot of popularity among other swarm intelligence methods due to its various characteristics like few tuning parameters, simplicity and easy to use, scalable, and most importantly its ability to provide faster convergence by maintaining the right balance between the exploration and exploitation during the search. Therefore, an enhanced GWO has been proposed in combination with SVM to determine the optimum subset of tumor features for accurate identification of benign and malignant tumor. The proposed approach has been tested and compared with numerous existing, state-of-the-art as well as recently published breast cancer classification approaches on the standard benchmark Wisconsin Diagnostic Breast Cancer (WDBC) database. The proposed approach outperforms all the compared approaches by improving the classification accuracy to 98.24% demonstrating its effectiveness in identifying the breast cancer.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant