Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Breast Cancer Stage Identification Using Machine Learning

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Breast Cancer Stage Identification Using Machine Learning

Similar Papers
  • PDF Download Icon
  • Research Article
  • Cite Count Icon 88
  • 10.1136/bmj-2022-073800
Development and internal-external validation of statistical and machine learning models for breast cancer prognostication: cohort study
  • May 10, 2023
  • BMJ
  • Ash Kieran Clift + 6 more

ObjectiveTo develop a clinically useful model that estimates the 10 year risk of breast cancer related mortality in women (self-reported female sex) with breast cancer of any stage, comparing results...

  • Research Article
  • Cite Count Icon 1
  • 10.1097/cej.0000000000000892
Application of machine learning in the analysis of multiparametric MRI data for the differentiation of treatment responses in breast cancer: retrospective study.
  • Jun 19, 2024
  • European journal of cancer prevention : the official journal of the European Cancer Prevention Organisation (ECP)
  • Jinhua Wang + 4 more

The objective of this study is to develop and validate a multiparametric MRI model employing machine learning to predict the effectiveness of treatment and the stage of breast cancer. The study encompassed 400 female patients diagnosed with breast cancer, with 200 individuals allocated to both the control and experimental groups, undergoing examinations in Shenzhen, China, during the period 2017-2023. This study pertains to retrospective research. Multiparametric MRI was employed to extract data concerning tumor size, blood flow, and metabolism. The model achieved high accuracy, predicting treatment outcomes with an accuracy of 92%, sensitivity of 88%, and specificity of 95%. The model effectively classified breast cancer stages: stage I, 38% ( P = 0.027); stage II, 72% ( P = 0.014); stage III, 50% ( P = 0.032); and stage IV, 45% ( P = 0.041). The developed model, utilizing multiparametric MRI and machine learning, exhibits high accuracy in predicting the effectiveness of treatment and breast cancer staging. These findings affirm the model's potential to enhance treatment strategies and personalize approaches for patients diagnosed with breast cancer. Our study presents an innovative approach to the diagnosis and treatment of breast cancer, integrating MRI data with machine learning algorithms. We demonstrate that the developed model exhibits high accuracy in predicting treatment efficacy and differentiating cancer stages. This underscores the importance of utilizing MRI and machine learning algorithms to enhance the diagnosis and individualization of treatment for this disease.

  • Research Article
  • Cite Count Icon 12
  • 10.1093/bib/bbae628
Comprehensive bioinformatics and machine learning analyses for breast cancer staging using TCGA dataset
  • Nov 22, 2024
  • Briefings in Bioinformatics
  • Saurav Chandra Das + 5 more

Breast cancer is an alarming global health concern, including a vast and varied set of illnesses with different molecular characteristics. The fusion of sophisticated computational methodologies with extensive biological datasets has emerged as an effective strategy for unravelling complex patterns in cancer oncology. This research delves into breast cancer staging, classification, and diagnosis by leveraging the comprehensive dataset provided by the The Cancer Genome Atlas (TCGA). By integrating advanced machine learning algorithms with bioinformatics analysis, it introduces a cutting-edge methodology for identifying complex molecular signatures associated with different subtypes and stages of breast cancer. This study utilizes TCGA gene expression data to detect and categorize breast cancer through the application of machine learning and systems biology techniques. Researchers identified differentially expressed genes in breast cancer and analyzed them using signaling pathways, protein–protein interactions, and regulatory networks to uncover potential therapeutic targets. The study also highlights the roles of specific proteins (MYH2, MYL1, MYL2, MYH7) and microRNAs (such as hsa-let-7d-5p) that are the potential biomarkers in cancer progression founded on several analyses. In terms of diagnostic accuracy for cancer staging, the random forest method achieved 97.19%, while the XGBoost algorithm attained 95.23%. Bioinformatics and machine learning meet in this study to find potential biomarkers that influence the progression of breast cancer. The combination of sophisticated analytical methods and extensive genomic datasets presents a promising path for expanding our understanding and enhancing clinical outcomes in identifying and categorizing this intricate illness.

  • Research Article
  • 10.37047/jos.galenos.2026.2025-9-2
A Survey on Machine and Deep Learning Techniques for Breast Cancer Prediction
  • Jan 22, 2026
  • Journal of Oncological Sciences
  • Pallavi Jadhav + 1 more

Breast cancer is a leading cause of cancer-related mortality among women worldwide, making early and accurate diagnosis essential for effective treatment and improved patient outcomes.In recent years, machine learning (ML) and deep learning (DL) techniques have emerged as promising tools for predicting and classifying breast cancer using gene expression and clinical data.However, existing studies face several limitations.Many rely solely on ML or DL approaches, lack comprehensive strategies for feature selection or extraction, and demonstrate inconsistent performance across datasets.These gaps result in models that are insufficiently accurate, uninterpretable, or unable to generalize well to unseen data.This work aims to address these challenges by conducting a detailed literature survey of existing ML and DL models applied to breast cancer prediction.The objectives include identifying common datasets, performance metrics, model types, and feature-engineering techniques.A structured methodology was followed to analyze peer-reviewed studies and extract trends in performance and limitations.Findings show that, while DL models outperform traditional ML in terms of accuracy, they often lack transparency and robust feature engineering.In conclusion, a unified approach combining advanced feature selection and extraction methods with DL techniques is necessary to develop accurate, generalizable breast cancer prediction systems.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 10
  • 10.3390/cancers16101864
Identification of Gene Expression in Different Stages of Breast Cancer with Machine Learning.
  • May 14, 2024
  • Cancers
  • Ali Abidalkareem + 4 more

Determining the tumor origin in humans is vital in clinical applications of molecular diagnostics. Metastatic cancer is usually a very aggressive disease with limited diagnostic procedures, despite the fact that many protocols have been evaluated for their effectiveness in prognostication. Research has shown that dysregulation in miRNAs (a class of non-coding, regulatory RNAs) is remarkably involved in oncogenic conditions. This research paper aims to develop a machine learning model that processes an array of miRNAs in 1097 metastatic tissue samples from patients who suffered from various stages of breast cancer. The suggested machine learning model is fed with miRNA quantitative read count data taken from The Cancer Genome Atlas Data Repository. Two main feature-selection techniques have been used, mainly Neighborhood Component Analysis and Minimum Redundancy Maximum Relevance, to identify the most discriminant and relevant miRNAs for their up-regulated and down-regulated states. These miRNAs are then validated as biological identifiers for each of the four cancer stages in breast tumors. Both machine learning algorithms yield performance scores that are significantly higher than the traditional fold-change approach, particularly in earlier stages of cancer, with Neighborhood Component Analysis and Minimum Redundancy Maximum Relevance achieving accuracy scores of up to 0.983 and 0.931, respectively, compared to 0.920 for the FC method. This study underscores the potential of advanced feature-selection methods in enhancing the accuracy of cancer stage identification, paving the way for improved diagnostic and therapeutic strategies in oncology.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 6
  • 10.21271/zjpas.34.2.3
Comprehensive Study for Breast Cancer Using Deep Learning and Traditional Machine Learning
  • Apr 12, 2022
  • ZANCO JOURNAL OF PURE AND APPLIED SCIENCES
  • Chiman Haydar Salh + 1 more

Comprehensive Study for Breast Cancer Using Deep Learning and Traditional Machine Learning

  • Research Article
  • Cite Count Icon 1
  • 10.1089/genbio.2022.29028.dko
A Data-Driven Lens to Understand Human Biology: An Interview with Daphne Koller
  • Jun 1, 2022
  • GEN Biotechnology
  • Daphne Koller + 1 more

A Data-Driven Lens to Understand Human Biology: An Interview with Daphne Koller

  • Book Chapter
  • Cite Count Icon 8
  • 10.1201/9781003133681-8
Advances in Machine Learning and Deep Learning Approaches for Mammographic Breast Density Measurement for Breast Cancer Risk Prediction: An Overview
  • Jul 7, 2021
  • Shivaji D Pawar + 2 more

Breast cancer and mammographic breast density are strongly associated with each other. As breast density increases, the chance of masking breast cancer also increases which reduces the sensitivity of measuring mammographic density. Mammographic breast density measurement is currently subjective with the support of BIRADS classification. These subjective qualitative assessments are prone to substantial inter-reader and intra-reader variations. Accurate and precise measurement of density is crucial for the correct prediction of breast cancer risk. Breast density measurement using machine learning and deep learning are gaining significant momentum due to the rapid developments in this field. The basic objective of this chapter is to provide a systematic overview of selected research articles on the machine and deep learning approaches, specifically to breast density measurement for cancer prediction. Comparative analysis of different recent articles is provided in terms of classification accuracy and computational complexity to decide the future path of research in this area. This chapter will be helpful for research scholars in this domain toward designing a novel algorithm for the measurement of breast density using either machine learning or deep learning.

  • Research Article
  • Cite Count Icon 2
  • 10.1016/j.clbc.2025.05.001
Ultrasound Radiomics-Based Machine Learning and SHapley Additive exPlanations Method Predicting Pathological Prognostic Stage in Breast Cancer: A Bicentric and Validation Study.
  • May 1, 2025
  • Clinical breast cancer
  • Lei Chen + 11 more

Ultrasound Radiomics-Based Machine Learning and SHapley Additive exPlanations Method Predicting Pathological Prognostic Stage in Breast Cancer: A Bicentric and Validation Study.

  • Research Article
  • Cite Count Icon 9
  • 10.1007/s11307-023-01823-8
Application of Machine Learning Analyses Using Clinical and [18F]-FDG-PET/CT Radiomic Characteristics to Predict Recurrence in Patients with Breast Cancer.
  • May 16, 2023
  • Molecular imaging and biology
  • Kodai Kawaji + 8 more

To develop and identify machine learning (ML) models using pretreatment clinical and 2-deoxy-2-[18F]fluoro-D-glucose positron emission tomography ([18F]-FDG-PET)-based radiomic characteristics to predict disease recurrences in patients with breast cancers who underwent surgery. This retrospective study included 112 patients with 118 breast cancer lesions who underwent [18F]-FDG-PET/ X-ray computed tomography (CT) preoperatively, and these lesions were assigned to training (n=95) and testing (n=23) cohorts. A total of 12 clinical and 40 [18F]-FDG-PET-based radiomic characteristics were used to predict recurrences using 7 different ML algorithms, namely, decision tree, random forest (RF), neural network, k-nearest neighbors, naive Bayes, logistic regression, and support vector machine (SVM) with a 10-fold cross-validation and synthetic minority over-sampling technique. Three different ML models were created using clinical characteristics (clinical ML models), radiomic characteristics (radiomic ML models), and both clinical and radiomic characteristics (combined ML models). Each ML model was constructed using the top ten characteristics ranked by the decrease in Gini impurity. The areas under ROC curves (AUCs) and accuracies were used to compare predictive performances. In training cohorts, all 7 ML algorithms except for logistic regression algorithm in the radiomics ML model (AUC = 0.760) achieved AUC values of >0.80 for predicting recurrences with clinical (range, 0.892-0.999), radiomic (range, 0.809-0.984), and combined (range, 0.897-0.999) ML models. In testing cohorts, the RF algorithm of combined ML model achieved the highest AUC and accuracy (95.7% (22/23)) with similar classification performance between training and testing cohorts (AUC: training cohort, 0.999; testing cohort, 0.992). The important characteristics for modeling process of this RF algorithm were radiomic GLZLM_ZLNU and AJCC stage. ML analyses using both clinical and [18F]-FDG-PET-based radiomic characteristics may be useful for predicting recurrence in patients with breast cancers who underwent surgery.

  • Conference Article
  • Cite Count Icon 5
  • 10.1109/ccet56606.2022.10080757
Analysis of Breast Cancer in Early Stage by UsingMachine LearningAlgorithms: A Review
  • Dec 23, 2022
  • Jasjeet Kaur Sandhu + 2 more

The goal of using Machine Learning (ML), a branch of Artificial Intelligence (AI) technology, is to increase the accuracy and efficiency of doctors' work. After skin cancer, breast cancer is the kind of cancer that affects women the most often. Breast cancer is caused by an unchecked multiplication of breast cells. According to the national cancer institute, 287,850 are the Estimated New Cases in 2022, with 15% of all new cancer cases. 43,250 will be the. Estimate deaths in 2022 in the USA. In the United States, female breast cancer accountsfor 15% of all newly diagnosed instances of cancer. An analysis of breast cancer data from 2018 in India revealed 1,62,468 fresh recorded cases and 87,090 deaths from the condition. According to the findings of the most recent research, the state of Kerala has the highest cancer incidence in all of India. This review paper will understand the various stages of Breast cancer, look at datasets and evaluate supervised and unsupervised Machine Learning (ML) algorithms to see which ones are best for early breast cancer prediction made possible by the use of machine learning or deep learning.

  • Conference Article
  • Cite Count Icon 3
  • 10.1109/icidca56705.2023.10100252
Meta Analysis of Human Body Diseases with the Application of Machine Learning
  • Mar 14, 2023
  • Nikhil Verma + 2 more

Machine learning in medical applications is one of the focus areas of the researchers these days. Machine Learning with the application of Artificial Intelligence is not only giving solutions to the complex problems but also revolutionised the medical field. The main motive of machine learning is to improve its learning process over time by taking all the relevant data and information in the form of different inputs and observations. This study reviews different medical disease prediction and detection techniques with the help of distinct deep learning & machine learning models. The problems related to medical diseases, like cancer related diseases, heart, lung, thyroid and kidney diseases are being discussed in this article. Detection and analysing of medical diseases is one of the prominent applications of machine and deep learning. Deep learning as a technology offers a huge set of different and innovative tools which are relevant to different issues faced in the field of medical image processing. This study will discuss about the applications of Machine Learning, and then discuss some of the advancements done in different diseases like breast cancer, heart disease, skin disease, kidney disease etc.

  • Research Article
  • 10.1158/1557-3265.aimachine-a051
Abstract A051: Multi-omic explainable machine learning improves cancer treatment outcome prediction
  • Jul 10, 2025
  • Clinical Cancer Research
  • Khoa A Tran + 15 more

Background: Advancements in multi-omics data integration and explainable Machine Learning (ML) have shown promise in precision oncology. Multi-omic data used to train ML models may include genomics, transcriptomics and histopathology to characterize cancer cells and the tumor microenvironment (TME). Explainability methods, such as SHAP, have enabled researchers and clinicians to unravel the decision-making rationale of ML models predicting cancer progression and treatment response. We developed an explainable ML framework that incorporates multi-omic features of cancer and the TME. This framework was applied to predict patient response to neoadjuvant chemotherapy (NAC) in breast cancer and immune checkpoint inhibitor (ICI) in melanoma. Methods: For breast cancer, we used the cohort from Sammut et al. [1] (n=157 training, n=75 test). For melanoma, we assembled a cohort comprising 229 patients (n=138 training, n=53 test cutaneous, n=38 test non-cutaneous) from five independent studies. We improved the performance of the ensemble ML models in Sammut et al. [1] by implementing a shared-learning architecture to enable component models to influence each other as training progresses. We applied this ensemble (Ens:LR+RF+SVM) to predict NAC response in breast cancer and ICI response in melanoma by integrating clinical, DNA sequencing, RNA sequencing and histopathology (only for breast cancer) data. For melanoma, we also trained three single ML models (LR, RF, and SVM) and another ensemble (Ens:LR+RF), and introduced a novel dual utility of SHAP for feature-selection during training and biomarker threshold identification during validation. Results: The ensemble model trained on the multi-omic breast cancer features achieved ROC-AUC of 0.88 and showed a potential 25% reduction in false positives (i.e., incorrect predictions of good response) compared to its predecessor from Sammut et al. [1]. In the melanoma cohort, the Ens:LR+RF model achieved ROC-AOC of 0.77 but was outperformed by the RF model, ROC-AUC 0.78. SHAP revealed unique interactions between each ML model and the feature space, resulting in distinct training feature sets per model. During validation, the intersection between feature values and SHAP scores revealed numerical thresholds underpinning good versus poor responses of clinically meaningful biomarkers such as neoantigen load (>2.25 good, <2.25 poor, values in log10 scale). Across these two studies, we developed and open-sourced a scalable and versatile ML workflow (xML-workFLow) for rapid experimentation in biomedical research. Conclusions: This work showcases the potential of multi-omics explainable ML in advancing precision oncology to improve treatment outcome prediction. With further experimental validation, the use of explainable ML to determine numerical thresholds could guide the development of companion diagnostics and inform combination therapeutic strategies.

  • Research Article
  • Cite Count Icon 1
  • 10.17762/ijritcc.v11i10s.7603
Early Breast Cancer Prediction using Machine Learning and Deep Learning Techniques
  • Oct 7, 2023
  • International Journal on Recent and Innovation Trends in Computing and Communication
  • Swati B Patil + 2 more

Breast Cancer (BC) is a considered as one of the utmost lethal diseases across the globe that has a very high morbidity and mortality rate. Accurate and early prediction along with diagnosis is one of the most crucial characteristics for the treatment of Breast Cancer. Doctors can have an edge over Breast cancer if they are able to predict it in its early stages using deep learning and machine learning techniques. This paper proposed consists of comparison between the and accuracy of various machine learning models like Support vector machine (SVM), K-Nearest Neighbours (KNN), Naïve Bayes (NB), Logistic Regression (LR), Random Forest (RF), Decision Tree (DT), XGB Classifier and deep learning model of Artificial neural networks (ANN) for the precise detection of breast cancer.
 The most crucial properties from the database have been chosen using one feature-selection technique. Correlation is also used to choose the most correlated features from the data. Implementing the ANN model consists of one input layer, two hidden layers, and one output layer. All Machine Learning models and ANN model are then applied to selected features. The results demonstrated that the SVM classifier achieved the highest performance with an accuracy of ~98.24%.

  • Research Article
  • 10.1093/clinchem/hvae106.483
B-122 Improving the Diagnostic Utility of Multiplexed Assays in Laboratory Medicine with Machine Learning
  • Oct 2, 2024
  • Clinical Chemistry
  • H A Miller + 3 more

Background The application of machine learning (ML) and artificial intelligence (AI) in precision medicine is of great interest and continues to grow. ML approaches offer a flexible and data-centric methodology for modeling diagnostic and prognostic trends, capable of identifying complex non-linear patterns and relationships between predictor variables and outcomes within large data sets. Although ML holds great promise, there are several common pitfalls that must be avoided for ML models to provide clinical utility. In the current study, we provide a rigorous and systematic evaluation of the utility of ML algorithms to predict disease in several representative data sets generated via multiplexed biomarker assays. We compare five major classes of ML algorithms compared to more traditional combinatorial analysis (logistic regression) to offer perspective against the current gold standard. Methods We obtained five data sets from published literature that include measurements of multiplexed biomarkers in patients with obstructive sleep apnea, sepsis, and breast and colon cancers in comparison to healthy controls. These datasets were generated by bead-based and planar multiplexed immunoassays. We chose representative ML algorithms from the classes of support vector machines (SVM), random forests (RF), neural networks (NNet), and extreme gradient boosting (xgBoost) in comparison to logistic regression (LR). Datasets were randomly split into internal validation and external validation sets via 10 iterations of permutation analysis. We tuned and validated the models using 5-fold cross validation, then tested those on external subsets. We report the average of several model classification metrics and associated variances across all resampled data sets. Results For the classification of sepsis using inflammatory cytokines measured on the Luminex bead-based immunoassay, RF outperformed LR in external validation sets by a 9.8% increase in AUROC. For classification of breast cancer patients, xgBoost and RF performed best with 20.0% and 16.9% increases in AUROC, respectively, whereas for classification of colon cancer patients, NNet outperformed LR by 11.5%. Across all datasets, on average, all ML algorithms outperformed LR in the internal validation sets by a 7.9% (6.3% - 9.9%) increase in AUROC. However, in external validation sets, only xgBoost and SVM outperformed LR, where RF and NNet showed signs of overfitting. Conclusions Overall, certain ML algorithms show improvement over LR for diagnostic applications using multiplexed assays, by up to 20% increase in AUROC. Interpretation of ML results must be performed rigorously to prevent overly optimistic and unrealistic conclusions. In the context of laboratory medicine, our results showcase the utility of ML in clinical applications, while highlighting potential disadvantages of ML. This analysis underscores the need for establishing rigorous guidelines to support data analytics during the development of novel multiplexed assays, which is needed to advance precision medicine.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant