Survival Period Prediction in Cervical Cancer Patients Using the Selective Stacking Technique.
This study aims to increase the effectiveness of cervical cancer treatment by developing a survival prediction model using an innovative ensemble machine learning approach, namely the selective stacking technique. Patient data obtained from the Faculty of Medicine, Chiang Mai University, Thailand, were utilized to validate the real-world applicability of the proposed approach. The selective stacking model employed a two-stage machine learning framework in which outputs from base machine learning models were systematically combined through meta-level learning. Importantly, the performance of the proposed model was compared with that reported in previous studies that relied on individual machine learning algorithms as baselines. To provide deeper insight into the predictive mechanisms of the model, local interpretable model-agnostic explanations were applied to assess feature importance and identify the most influential factors contributing to model predictions. The classification model developed using the selective stacking technique demonstrated a marked improvement in prediction accuracy, achieving an accuracy of 91.41%. The regression model also showed robust performance, with a root mean square error of 18.92 and an r value of 0.669. Feature importance analysis indicated that side effect status involving surrounding organs emerged as the most influential factor in survival prediction. The selective stacking model exhibited superior predictive performance compared with the base models, suggesting that this approach offers a promising strategy for cervical cancer survival prediction and may support the development of more personalized treatment planning.
- Research Article
25
- 10.1002/cam4.5477
- Dec 8, 2022
- Cancer Medicine
BackgroundPrediction models with high accuracy rates for nonmetastatic cervical cancer (CC) patients are limited. This study aimed to construct and compare predictive models on the basis of machine learning (ML) algorithms for predicting the 5‐year survival status of CC patients through using the Surveillance, Epidemiology, and End Results public database of the National Cancer Institute.MethodsThe data registered from 2004 to 2016 were extracted and randomly divided into training and validation cohorts (8:2). The least absolute shrinkage and selection operator (LASSO) regression was employed to identify significant factors. Then, four predictive models were constructed, including logistic regression (LR), random forest (RF), support vector machine (SVM), and extreme gradient boosting (XGBoost). The predictive models were evaluated and compared using Receiver‐operating characteristics with areas under the curves (AUCs) and decision curve analysis (DCA), respectively.ResultsA total of 13,802 patients were involved and classified into training (N = 11,041) and validation (N = 2761) cohorts. By using the LASSO regression method, seven factors were identified. In the training cohort, the XGBoost model showed the best performance (AUC = 0.8400) compared to the other three models (all p < 0.05 by Delong's test). In the validation cohort, the XGBoost model also demonstrated a superior prediction ability (AUC = 0.8365) than LR and SVM models (both p < 0.05 by Delong's test), although the difference was not statistically significant between the XGBoost and the RF models (p = 0.4251 by Delong's test). Based on the DCA results, the XGBoost model was also superior, and feature importance analysis indicated that the tumor stage was the most important variable among the seven factors.ConclusionsThe XGBoost model proved to be an effective algorithm with better prediction abilities. This model is proposed to support better decision‐making for nonmetastatic CC patients in the future.
- Research Article
21
- 10.1016/j.ijhydene.2024.11.329
- Nov 28, 2024
- International Journal of Hydrogen Energy
Improving syngas yield and quality from biomass/coal co-gasification using cooperative game theory and local interpretable model-agnostic explanations
- Research Article
13
- 10.7717/peerj-cs.1916
- May 17, 2024
- PeerJ Computer Science
Background Cancer is positioned as a major disease, particularly for middle-aged people, which remains a global concern that can develop in the form of abnormal growth of body cells at any place in the human body. Cervical cancer, often known as cervix cancer, is cancer present in the female cervix. In the area where the endocervix (upper two-thirds of the cervix) and ectocervix (lower third of the cervix) meet, the majority of cervical cancers begin. Despite an influx of people entering the healthcare industry, the demand for machine learning (ML) specialists has recently outpaced the supply. To close the gap, user-friendly applications, such as H2O, have made significant progress these days. However, traditional ML techniques handle each stage of the process separately; whereas H2O AutoML can automate a major portion of the ML workflow, such as automatic training and tuning of multiple models within a user-defined timeframe. Methods Thus, novel H2O AutoML with local interpretable model-agnostic explanations (LIME) techniques have been proposed in this research work that enhance the predictability of an ML model in a user-defined timeframe. We herein collected the cervical cancer dataset from the freely available Kaggle repository for our research work. The Stacked Ensembles approach, on the other hand, will automatically train H2O models to create a highly predictive ensemble model that will outperform the AutoML Leaderboard in most instances. The novelty of this research is aimed at training the best model using the AutoML technique that helps in reducing the human effort over traditional ML techniques in less amount of time. Additionally, LIME has been implemented over the H2O AutoML model, to uncover black boxes and to explain every individual prediction in our model. We have evaluated our model performance using the findprediction() function on three different idx values (i.e., 100, 120, and 150) to find the prediction probabilities of two classes for each feature. These experiments have been done in Lenovo core i7 NVidia GeForce 860M GPU laptop in Windows 10 operating system using Python 3.8.3 software on Jupyter 6.4.3 platform. Results The proposed model resulted in the prediction probabilities depending on the features as 87%, 95%, and 87% for class ‘0’ and 13%, 5%, and 13% for class ‘1’ when idx_value=100, 120, and 150 for the first case; 100% for class ‘0’ and 0% for class ‘1’, when idx_value= 10, 12, and 15 respectively. Additionally, a comparative analysis has been drawn where our proposed model outperforms previous results found in cervical cancer research.
- Research Article
1
- 10.26483/ijarcs.v15i5.7126
- Oct 20, 2024
- international journal of advanced research in computer science
Cervical cancer is a highly prevalent malignancy affecting women worldwide, ranking as the seventh most common cancer globally. This study aims to systematically review and analyze cervical cancer survival predictions using machine learning (ML) algorithms. A comprehensive search was conducted across Scopus and PubMed databases in February 2024. Extracted articles were screened using Hubmeta software, with duplicates and non-relevant studies excluded. The final selection, comprising 24 articles, focused on survival predictions through ML techniques. These studies, published mostly post-2019, included datasets ranging from 75 to 9,462 cervical cancer patients and up to 91,294 squamous cell samples. The most commonly applied ML models were Random Forest (RF), Neural Networks (NN), Support Vector Machines (SVM), Ensemble and Hybrid Learning, and Deep Learning (DL). The area under the curve (AUC) for these models ranged from 0.84 to 0.9875, demonstrating their strong predictive capabilities. Clinical patient records were the primary data source. Meta-analysis was performed on the extracted data using GraphPad Prism for descriptive statistics and One-Way ANOVA. No significant differences were found between group means, as evidenced by an R-squared value of 0.1459. This result indicates that the independent variable (year of study) explained only 14.59% of the variance in ML model performance. The study found that the use of ML models has increased over time, particularly with Convolutional Neural Networks (CNNs) such as the ResNet50 model, which demonstrated superior accuracy metrics, including over 90% accuracy for the ResNet152 variant. These findings suggest that integrating multi-dimensional data with ML models holds significant potential for improving survival predictions in cervical cancer patients. Future research is recommended to develop tailored ML algorithms with even higher predictive accuracy for cervical cancer survival.
- Research Article
1
- 10.21037/tcr-2024-2304
- May 1, 2025
- Translational cancer research
Cervical cancer (CC) is one of the most common gynecological malignancies. Previous studies have shown that the prognosis of CC is affected by many factors. Our study aimed to identify key prognostic factors and use machine learning and deep learning algorithms to construct models to predict the overall survival (OS) of CC patients. Data of CC patients collected between 2007 and 2016 were collected from the Surveillance, Epidemiology, and End Results (SEER) database, and were randomly divided into the training set (1,743 patients) and test set (747 patients). Moreover, in order to enhance the practical application of the model, we conducted an X-tile analysis to categorize the patients into three distinct strata based on their age and tumor size. Least absolute shrinkage and selection operator (LASSO) and multivariate Cox regression were performed to identify the independent prognostic factors for OS, which were further used to construct CoxBoost, RandomForest, SuperPC XGBoost, and DeepSurv survival models to predict 1-, 3-, and 5-year OS. The parameters, including age, marital status, grade, tumor size, surgery, radiation, race, the American Joint Committee on Cancer (AJCC)_stage, AJCC_T, and AJCC_M, were associated with survival and were further incorporated into the five models. The concordance index (C-index) value was 0.858, 0.848, 0.849, 0.840, and 0.869, respectively, and the receiver operating characteristic (ROC) curves showed exceptional predictive performance. Among the five models, DeepSurv was the model with best performance. The ROC curve validated the area under the curve (AUC) values for 1-year OS, 3-year OS, and 5-year OS, which were 0.936, 0.915, and 0.900, respectively. The prognostic model conducted by DeepSurv algorithm and the independent prognostic factors can potentially be applied in making personalized treatment plans and evaluating the prognosis of CC patients.
- Research Article
1
- 10.7717/peerj-cs.3258
- Oct 28, 2025
- PeerJ Computer Science
The increasing sophistication of evolving malware types and attack techniques has rendered traditional antivirus solutions inadequate, particularly in mitigating zero-day threats. To address this challenge, Machine Learning (ML) and Deep Learning (DL)-based approaches have been developed, demonstrating significant efficacy and high accuracy in malware classification. However, the black box nature of these models raises significant concerns in terms of transparency and interpretability. This study presents a comprehensive evaluation of Ensemble Learning and Deep Learning methods for static analysis-based malware classification, which allows joint analysis of Application Programming Interface (API) calls and Dynamic Link Library (DLL) data. In the study, a specially designed Convolutional Neural Network (CNN)-Gated Recurrent Units (GRU)-3 model is trained using a tailored dataset consisting of malicious and secure software. In order to better understand the model’s performance, feature importance analysis was performed using SHapley additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME) Explainable Artificial Intelligence (XAI) techniques and the reliability of model decisions was increased. The proposed model was compared with DL models such as CNN, Long Short-Term Memory (LSTM), and GRU, as well as traditional ML algorithms such as Extreme Gradient Boosting (XGB), Extra Trees Classifier (ETC), K-Nearest Neighbor (KNN), and Random Forest (RF). While the traditional XGB model achieved the highest overall performance with a 99.81% accuracy rate, our proposed CNN-GRU-3 model demonstrated the best performance among all DL models, reaching a 99.37% accuracy rate. This study presents a powerful framework that provides both high accuracy in malware detection and makes the decision mechanism more transparent.
- Research Article
13
- 10.1080/01443615.2019.1652888
- Oct 12, 2019
- Journal of Obstetrics and Gynaecology
The purpose of this study was to develop and validate a nomogram for individual prediction of recurrence and disease-free survival (DFS) among lymph node (LN)-negative early-stage (I–IIA) cervical cancer (CC) patients treated with Type B or Type C2 hysterectomy. Data were collected from patients diagnosed with CC between 1995 and 2017 at the Gynecological Oncology Department, Tepecik Training and Research Hospital. A total of 194 cases with stage IA2–IIA CC were evaluated retrospectively. Patients with stage IA2–IIA CC who underwent radical (Type C2) or modified radical (Type B) hysterectomy and pelvic ± paraaortic LN dissection with LN negativity were included in the study. The relationships between prognostic factors such as stage, tumour size, parametrial involvement, vaginal cuff margin, endomyometrial infiltration, and lymphovascular space invasion status and DFS were compared using a univariable Cox regression model. When the nomogram was prepared, the scores of the risk factors were collected, and we observed that scores were at least 0 to a maximum of 414 points. The concordance-index for the nomogram was 0.895 (95% confidence interval, 0.79–0.99). The nomogram based on the indicated prognostic factors yielded excellent results in predicting recurrence in early-stage CC patients without LN metastasis who underwent radical hysterectomy.Impact statementWhat is already known on this subject? Pathology of radical hysterectomy specimens in patients with early-stage cervical cancer provides information that has predictive prognostic potential. In addition to FIGO stage, other important prognostic factors are lymph node status, tumour size, parametrial involvement, vaginal cuff margin status, endomyometrial infiltration, histological type, patient age, lymphovascular space invasion, histological grade, and depth of cervical stromal invasion.What do the results of this study add? In this study, patients with early-stage cervical cancer who underwent radical and modified radical hysterectomy without retroperitoneal lymph node involvement were evaluated, and recurrence development and factors affecting disease-free survival were investigated. A nomogram consisting of factors influencing disease-free survival was constructed. The total score was determined according to the status of all risk factors. This allowed clear definition of the risk for each patient. A nomogram predicting recurrence in patients with stages IA2–IIA cervical cancer with radical hysterectomy without lymph node involvement has not previously been published.What are the implications of these findings for clinical practice and/or further research? Our study investigated early-stage cervical cancer (CC) patients without lymph node (LN) metastasis. Cox regression analysis was performed with six prognostic factors: FIGO stage, tumour size, parametrial margin infiltration, vaginal cuff margin involvement, endomyometrial infiltration, and LVSI positivity. The nomogram was constructed based on the results of Cox regression. The C-index for the nomogram was 0.895 (95% CI, 0.79–0.99). These results can be considered excellent. The higher concordance index in our study indicates that these six factors may be more valuable in predicting recurrence development in CC patients.
- Research Article
10
- 10.57197/jdr-2024-0081
- Jan 1, 2024
- Journal of Disability Research
This research introduces a novel approach, termed “explainable federated learning,” designed for privacy-preserving autism prediction in toddlers using deep learning (DL) techniques. The primary objective is to contribute to the development of efficient screening methods for autism spectrum disorder (ASD) while safeguarding individual privacy. The methodology encompasses multiple stages, starting with exploratory data analysis and progressing through machine learning (ML) algorithms, federated learning (FL), and model explainability using local interpretable model-agnostic explanations (LIME). Leveraging non-linear predictive models such as autoencoders, k-nearest neighbors, and multi-layer perceptron, this approach ensures accurate ASD predictions. The FL paradigm facilitates collaboration among multiple clients without centralizing raw data, addressing privacy concerns in medical data sharing. Privacy-preserving strategies, including differential privacy, are integrated to enhance data security. Furthermore, model explainability is achieved through LIME, providing interpretable insights into the prediction process. The experimental results demonstrate significant improvements in predictive accuracy and model interpretability compared to traditional ML approaches. Specifically, our approach achieved an average accuracy increase of 8% across all classifiers tested, demonstrating superior performance in both privacy and predictive metrics over traditional methods. The findings highlight the efficacy of the proposed methodology in advancing ASD screening methodologies in the era of DL applications.
- Research Article
8
- 10.1016/j.omto.2020.09.001
- Sep 5, 2020
- Molecular Therapy Oncolytics
A 70-Gene Signature for Predicting Treatment Outcome in Advanced-Stage Cervical Cancer
- Research Article
1
- 10.1038/s41598-025-23593-9
- Nov 18, 2025
- Scientific reports
Cervical cancer, predominantly caused by Human Papillomavirus (HPV) infection, remains a significant global health burden for women, contributing to elevated morbidity and mortality rates. Early and accurate prediction is critical in improving patient outcomes and optimizing healthcare resource allocation. While machine learning (ML) and deep learning (DL) methods-such as support vector machines, random forests, and convolutional neural networks-have demonstrated promise in disease prediction, model interpretability, computational efficiency, and rely on large, labeled datasets. Additionally, conventional diagnostic methods like piezoresistive, piezoelectric, and optical lever techniques are often cost-prohibitive and complex, limiting widespread use. This study proposes a hybrid ML framework that integrates H2O AutoML with an autoencoder-based feature extraction and Fisher Score-based feature selection. To enhance model transparency and clinical trust, Local Interpretable Model-Agnostic Explanations (LIME) and SHAP (SHapley Additive exPlanations) are employed. The workflow initiates with exploratory data analysis (EDA) and dimensionality reduction using a stacked autoencoder, followed by selection of the top predictive features via Fisher Score. The refined feature set is used to train multiple models via H2O AutoML, with the best-performing deep learning model selected. On the training dataset, the selected model achieved 95.24% accuracy, an AUC of 98.10, and a log loss of 0.1747. Cross-validation confirms the model's robustness with consistent AUC and log loss values. At the optimal F1 threshold of 0.517, the confusion matrix indicates an error rate of 5.75% for actual negatives and 2.59% for actual positives, leading to an overall error rate of 4.14%. LIME and SHAP are used to interpret predictions at the instance level, providing actionable insights for clinicians. These results demonstrate the effectiveness of combining AutoML with explainable AI and advanced feature engineering to enhance the predictive power and interpretability of cervical cancer risk models, offering a scalable solution for clinical decision support.
- Conference Article
- 10.54941/ahfe1006203
- Jan 1, 2025
- AHFE international
Diabetes mellitus is a global health concern affecting millions worldwide, with profound medical and socioeconomic implications. The increasing adoption of machine learning (ML) in healthcare has revolutionized clinical decision-making by enabling predictive diagnostics, personalized treatment plans, and efficient resource allocation. Despite their potential, many ML models are often regarded as "black boxes" due to their lack of transparency, which raises significant challenges in critical fields like healthcare, where explainability is crucial for ethical and accountable decision-making (Hassija et al., 2024).Explainable Artificial Intelligence (XAI) has emerged as a solution to address these challenges by making ML models more interpretable and fostering trust among healthcare practitioners and patients. This paper explores the integration of XAI techniques with ML models for diabetes prediction, emphasizing their potential to enhance transparency, trust, and clinical utility. We present a comparative analysis of popular XAI methods, such as SHAP (Shapley Additive Explanations), LIME (Local Interpretable Model-agnostic Explanations), and attention mechanisms, within the context of healthcare decision support. These techniques are evaluated based on interpretability, computational efficiency, and clinical applicability, highlighting the trade-offs between accuracy and transparency.The study underscores the critical role of interpretability in advancing trust and adoption of AI-driven solutions in healthcare, while addressing challenges such as balancing model performance with explainability. Finally, future directions for deploying explainable ML in healthcare are outlined, aiming to ensure ethical, transparent, and effective AI implementation.
- Research Article
25
- 10.1046/j.1525-1438.2000.00013.x
- Mar 3, 2000
- International Journal of Gynecological Cancer
The aim of this study was to determine the value of the measurement of serum VEGF and TGF-beta1 levels in the diagnosis of cervical cancer and to see whether these levels decrease after treatment for cervical cancer. We measured serum VEGF and TGF-beta1 levels through EIA in patients with CIN (n = 35), and cervical squamous cell cancer (n = 48). We also measured serum VEGF, TGF-beta1, and SCC antigen levels before and after radiotherapy in 13 cervical squamous cell cancer patients. The sizes of the tumors in those patients were measured by a computer tomography scan or magnetic resonance imaging. The serum VEGF levels were different between CIN and cervical cancer groups (P < 0.1), and the serum TGF-beta 1 levels in the cervical cancer group were lower than those in the other groups (P < 0.05). The serum VEGF levels were significantly related to the serum TGF-beta 1 levels in the cervical cancer patients (P < 0.01). In the cervical cancer patients, the decrease in the circulating VEGF levels after receiving radiotherapy was related to the decrease in tumor size (P < 0.01). While the measurement of serum VEGF level is adjuvant in diagnosing cervical cancers, serial serum VEGF level measurements may find a clinical use in the follow-up of women treated for cervical cancer.
- Research Article
10
- 10.1177/20552076251347785
- May 1, 2025
- Digital health
Heart failure (HF) is a primary contributor to morbidity and mortality among patients in intensive care units (ICUs), particularly those experiencing chronic critical illness (CCI). This study aims to develop and validate a machine learning (ML) model for predicting in-hospital mortality in CCI patients with HF. Retrospective data from over 200 hospitals were sourced from the Medical Information Mart for Intensive Care III (MIMIC-III), MIMIC-IV, and the eICU Collaborative Research Database (eICU-CRD). Only patients diagnosed with both CCI and HF were included. The MIMIC datasets served as the derivation cohort, while the eICU-CRD dataset was used for external validation. Key predictive variables were identified through recursive feature elimination. A range of ML algorithms, including random forest, K-nearest neighbors, and support vector machine (SVM), were evaluated alongside four other models. Model performance was assessed using the area under the receiver operating characteristic curve (AUROC). Model interpretability was enhanced through SHapley Additive exPlanations (SHAP) and local interpretable model-agnostic explanations. A total of 780 and 610 patients with CCI and HF were assigned to the derivation and validation cohorts, respectively. Eleven features were selected for model development. The SVM model demonstrated substantial predictive accuracy, with AUROC values of 0.781 and 0.675 in the derivation and validation cohorts. Feature importance analysis using SHAP identified Sequential Organ Failure Assessment score, oxyhemoglobin saturation, and blood pressure as key predictors. The SVM model developed reliably predicts in-hospital mortality in patients with CCI and HF, offering a valuable tool for early intervention and enhanced patient management.
- Conference Article
- 10.1109/icacrs67045.2025.11324170
- Dec 10, 2025
Explainable Artificial Intelligence (XAI) based tools such as SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) are extensively used in various detection and prediction approaches. These tools extract feature importance from the datasets and explain the contribution of the features (feature importance) towards detection /prediction output both locally and globally. In the current study a performance analysis is represented on the behaviour of LIME and SHAP explainability towards Denial-of-Service Attack detection in Internet of Things. There are numerous Black-box models including Machine Learning which show high detection accuracies in such case but the output is not interpretable by the security analyst most of the time. this drawback is overcome by introducing LIME and SHAP interpretability to the output of BlackBox model by analysing feature importance of the attack dataset towards detection accuracy. However, LIME and SHAPE has different behaviour towards model-interpretability. SHAP is powerful in global explanation where LIME works efficiently on local interpretation. We have shown that these two different tools perform on same detection accuracies of DoS attack using Machine learning model. A random forest classifier is first selected with high detection accuracy on a simulated DoS attack dataset and at the output SHAP and LIME are executed for achieving both local and global explainability. The comparison shows how SHAP and LIME show strength and weakness in explaining model’s behaviour both locally and globally.
- Abstract
- 10.1016/j.ijrobp.2013.06.087
- Sep 20, 2013
- International Journal of Radiation Oncology*Biology*Physics
The Use of 18FDG-PET Standard Uptake Value as a Metabolic Predictor of Bone Marrow Response to Radiation: Impact on Acute and Late Hematological Toxicity in Cervical Cancer Patients Treated With Chemoradiation Therapy