Improving Anomaly Detection in the HDFS Dataset with Novel Machine Learning Models and Techniques
This study addresses the limitations of existing anomaly detection methods on the HDFS dataset by introducing novel machine learning approaches, enhanced feature extraction, and techniques like SMOTE and temporal features, resulting in statistically significant improvements and setting new performance benchmarks.
With the growing scale and complexity of log data, manual anomaly detection has become increasingly time-consuming and error-prone, necessitating the development of robust machine learning-based solutions. The HDFS (Hadoop Distributed File System) dataset, a large-scale real-world log collection, serves as a standard benchmark for evaluating both supervised and unsupervised anomaly detection methods. Available in two variants - a reduced and a complete version - this dataset facilitates comprehensive performance comparisons. Our empirical analysis reveals significant limitations in existing approaches: supervised methods exhibit poor performance on the reduced dataset, while most unsupervised techniques underperform across both versions. To address these shortcomings, we introduce several novel machine learning approaches for log-based anomaly detection. Additionally, we investigate the effects of alternative feature extraction techniques. We also examine the application of Synthetic Minority Over-sampling Technique (SMOTE) to mitigate class imbalance in supervised learning, as well as the incorporation of temporal features encoding inter-log time intervals. Our experimental results demonstrate that the proposed methods achieve statistically significant improvements in detection accuracy over existing approaches on both HDFS dataset variants, establishing new benchmarks for log-based anomaly detection.
- Research Article
659
- 10.1109/access.2021.3083060
- Jan 1, 2021
- IEEE Access
Anomaly detection has been used for decades to identify and extract anomalous components from data. Many techniques have been used to detect anomalies. One of the increasingly significant techniques is Machine Learning (ML), which plays an important role in this area. In this research paper, we conduct a Systematic Literature Review (SLR) which analyzes ML models that detect anomalies in their application. Our review analyzes the models from four perspectives; the applications of anomaly detection, ML techniques, performance metrics for ML models, and the classification of anomaly detection. In our review, we have identified 290 research articles, written from 2000-2020, that discuss ML techniques for anomaly detection. After analyzing the selected research articles, we present 43 different applications of anomaly detection found in the selected research articles. Moreover, we identify 29 distinct ML models used in the identification of anomalies. Finally, we present 22 different datasets that are applied in experiments on anomaly detection, as well as many other general datasets. In addition, we observe that unsupervised anomaly detection has been adopted by researchers more than other classification anomaly detection systems. Detection of anomalies using ML models is a promising area of research, and there are a lot of ML models that have been implemented by researchers. Therefore, we provide researchers with recommendations and guidelines based on this review.
- Conference Article
5
- 10.1117/12.2552838
- Mar 20, 2020
In the semiconductor manufacturing industry, Automatic Defect Classification (ADC) plays an important role in maintaining high wafer inspection quality and reducing yield loss. ADC performance has benefitted from using machine learning (ML) algorithms; however, performance is negatively affected by the data imbalance and limited amounts of training data. Synthetic Minority Oversampling Technique (SMOTE) is an oversampling technique to adjust the skewed class distribution of a dataset so that the bias of the majority class is reduced. This paper shows that applying SMOTE achieved higher accuracy and purity on two imbalanced datasets, consisting of scanning electron microscopy (SEM) images collected with ASML-HMI eP™ and eScan® series inspection tools. The ML models are also less sensitive to the selection of hyperparameters when SMOTE is applied. We also show that better classification results can be obtained with less training samples with SMOTE; we conducted an experiment where a ML model trained on only 25% of samples with SMOTE achieved a higher ADC accuracy and purity performance compared to the same ML model trained on all samples but without SMOTE. In another experiment using a highly imbalanced SEM dataset with very few counts of the defect-of- interest (DOI), the combination of SMOTE and random undersampling of the majority class improves the accuracy by up to 5x while maintaining the same level of purity.
- Research Article
325
- 10.1109/access.2021.3053763
- Jan 1, 2021
- IEEE Access
Chronic Kidney Disease is one of the most critical illness nowadays and proper diagnosis is required as soon as possible. Machine learning technique has become reliable for medical treatment. With the help of a machine learning classifier algorithms, the doctor can detect the disease on time. For this perspective, Chronic Kidney Disease prediction has been discussed in this article. Chronic Kidney Disease dataset has been taken from the UCI repository. Seven classifier algorithms have been applied in this research such as artificial neural network, C5.0, Chi-square Automatic interaction detector, logistic regression, linear support vector machine with penalty L1 & with penalty L2 and random tree. The important feature selection technique was also applied to the dataset. For each classifier, the results have been computed based on (i) full features, (ii) correlation-based feature selection, (iii) Wrapper method feature selection, (iv) Least absolute shrinkage and selection operator regression, (v) synthetic minority over-sampling technique with least absolute shrinkage and selection operator regression selected features, (vi) synthetic minority over-sampling technique with full features. From the results, it is marked that LSVM with penalty L2 is giving the highest accuracy of 98.86% in synthetic minority over-sampling technique with full features. Along with accuracy, precision, recall, F-measure, area under the curve and GINI coefficient have been computed and compared results of various algorithms have been shown in the graph. Least absolute shrinkage and selection operator regression selected features with synthetic minority over-sampling technique gave the best after synthetic minority over-sampling technique with full features. In the synthetic minority over-sampling technique with least absolute shrinkage and selection operator selected features, again linear support vector machine gave the highest accuracy of 98.46%. Along with machine learning models one deep neural network has been applied on the same dataset and it has been noted that deep neural network achieved the highest accuracy of 99.6%.
- Research Article
73
- 10.1093/comjnl/bxaa056
- Jun 17, 2020
- The Computer Journal
The area of corporate bankruptcy prediction attains high economic importance, as it affects many stakeholders. The prediction of corporate bankruptcy has been extensively studied in economics, accounting and decision sciences over the past two decades. The corporate bankruptcy prediction has been a matter of talk among academic literature and professional researchers throughout the world. Different traditional approaches were suggested based on hypothesis testing and statistical modeling. Therefore, the primary purpose of the research is to come up with a model that can estimate the probability of corporate bankruptcy by evaluating its occurrence of failure using different machine learning models. As the dataset was not well prepared and contains missing values, various data mining and data pre-processing techniques were utilized for data preparation. Within this research, the task of resolving the issues induced by the imbalance between the two classes is approached by applying different data balancing techniques. We address the problem of imbalanced data with the random undersampling and Synthetic Minority Over Sampling Technique (SMOTE). We used five machine learning models (support vector machine, J48 decision tree, Logistic model tree, random forest and decision forest) to predict corporate bankruptcy earlier to the occurrence. We use data from 2009 to 2013 on Poland manufacturing corporates and selected the 64 financial indicators to be broken down. The main finding of the study is a significant improvement in predictive accuracy using machine learning techniques. We also include other economic indicators ratios, along with Altman’s Z-score variables related to profitability, liquidity, leverage and solvency (short/long term) to propose an efficient model. Machine learning models give better results while balancing the data through SMOTE as compared to random undersampling. The machine learning technique related to decision forest led to 99% accuracy, whereas support vector machine (SVM), J48 decision tree, Logistic Model Tree (LMT) and Random Forest (RF) led to 92%, 92.3%, 93.8% and 98.7% accuracy, respectively, with all predictive financial indicators. We find that the decision forest outperforms the other techniques and previous techniques discussed in the literature. The proposed method is also deployed on the web to assist regulators, investors, creditors and scholars to predict corporate bankruptcy.
- Research Article
9
- 10.1007/s11307-023-01823-8
- May 16, 2023
- Molecular imaging and biology
To develop and identify machine learning (ML) models using pretreatment clinical and 2-deoxy-2-[18F]fluoro-D-glucose positron emission tomography ([18F]-FDG-PET)-based radiomic characteristics to predict disease recurrences in patients with breast cancers who underwent surgery. This retrospective study included 112 patients with 118 breast cancer lesions who underwent [18F]-FDG-PET/ X-ray computed tomography (CT) preoperatively, and these lesions were assigned to training (n=95) and testing (n=23) cohorts. A total of 12 clinical and 40 [18F]-FDG-PET-based radiomic characteristics were used to predict recurrences using 7 different ML algorithms, namely, decision tree, random forest (RF), neural network, k-nearest neighbors, naive Bayes, logistic regression, and support vector machine (SVM) with a 10-fold cross-validation and synthetic minority over-sampling technique. Three different ML models were created using clinical characteristics (clinical ML models), radiomic characteristics (radiomic ML models), and both clinical and radiomic characteristics (combined ML models). Each ML model was constructed using the top ten characteristics ranked by the decrease in Gini impurity. The areas under ROC curves (AUCs) and accuracies were used to compare predictive performances. In training cohorts, all 7 ML algorithms except for logistic regression algorithm in the radiomics ML model (AUC = 0.760) achieved AUC values of >0.80 for predicting recurrences with clinical (range, 0.892-0.999), radiomic (range, 0.809-0.984), and combined (range, 0.897-0.999) ML models. In testing cohorts, the RF algorithm of combined ML model achieved the highest AUC and accuracy (95.7% (22/23)) with similar classification performance between training and testing cohorts (AUC: training cohort, 0.999; testing cohort, 0.992). The important characteristics for modeling process of this RF algorithm were radiomic GLZLM_ZLNU and AJCC stage. ML analyses using both clinical and [18F]-FDG-PET-based radiomic characteristics may be useful for predicting recurrence in patients with breast cancers who underwent surgery.
- Dissertation
- 10.17918/00001036
- Jun 1, 2020
Health-related behavior change is the core premise of behavioral medicine research and intervention yet remains unattainable and/or unsustainable for the majority of patients. Significant systemic, environmental and personal barriers limit access to biopsychoeducational information, medical providers and evidence-based treatments that have the power to motivate health promoting behaviors and reduce risk for lifetime development of chronic illness, such as diabetes and hypertension. Unfortunately, the impact of the field of health psychology is often methodologically limited by underpowered, unrepresentative samples and comparatively less powerful statistical methods/models for analyzing the complexities of chronic illness development and prognosis. However, this sphere of influence can easily expand through integration of advances in big data analytics and the application of more robust, predictively powerful machine learning (ML) models to traditional behavioral medicine research objectives. The process of improving health and wellbeing en masse begins with knowledge acquisition through dissemination of personally relevant, data driven health-related information to the general public. The present project sought to develop a comprehensive, multifactorial correlational model that classifies and predicts risk for lifetime development of three chronic diseases known to be responsive to health-related behavior change, diabetes, hypertension, and heart disease/heart attack, based on the intersection of demographic, health behavior, and healthrelated quality of life variables. Three ML models were utilized, in conjunction with Synthetic Minority Over-sampling Technique - Nominal Continuous (SMOTE - NC) to address imbalanced diagnostic groups: a) logistic regression, b) decision tree, and c) random forest. Experiments in dimensionality reduction including Factor Analysis of Mixed Data (FAMD) and permutation feature importance were also performed to enhance model efficiency and performance. Results supported the application of ML techniques to the classification of chronic illness, achieving target accuracy above 90% for two of the three diagnostic outcomes modeled: diabetes (92.1%) and heart disease/attack (94.4%). Hypertension (77.7%) was the only outcome that did not achieve target accuracy, which may be due in part to the relative complexity of risk and resiliency for this disease. Additionally, secondary analyses supported the classification of U.S. Veterans as a specialized subpopulation characterized by chronic illness disparity, particularly in regard to heart disease/attack. Future iterations of this project conducted with more expansive, feature rich data sources and optimizations in accuracy may allow for the eventual development of a web-based application with the ability to generate a personalized Chronic Illness Risk and Resiliency Profile for users that will include an adjusted predicted risk, a reduction in chronic illness risk which can be achieved through a combination of recommended, user specific health-related behavior changes. This research is designed to increase awareness of the role of behavior in health in order to empower the general public and motivate feasible increases in health-promoting behaviors across the United States.
- Research Article
1
- 10.1149/ma2025-01592799mtgabs
- Jul 11, 2025
- Electrochemical Society Meeting Abstracts
This study investigates advancements in machine learning (ML) techniques that improve the sensitivity and precision of biological sensing technologies, which are vital in fields such as medical diagnostics, environmental monitoring, and biotechnology. A systematic review was conducted using major academic databases, including PubMed, IEEE Xplore, Scopus, and Web of Science, focusing on research published between 2015 and 2024. A total of 75 peer-reviewed articles were analyzed to evaluate the impact of ML algorithms, such as deep learning, support vector machines, random forests, and ensemble methods on the functionality of biological sensors. The findings reveal that ML-driven approaches have significantly enhanced the performance of biosensors, wearable diagnostic platforms, and real-time environmental monitoring systems. These advancements include improvements in detection accuracy, signal-to-noise ratios, and real-time adaptability, enabling applications like early disease detection and precise environmental analysis. Furthermore, ML has optimized signal processing and facilitated the analysis of large and complex biological data sets, thereby increasing the overall efficiency of sensing technologies. Despite these achievements, several challenges persist. These include limited access to high-quality, annotated datasets, computational resource constraints, and issues with the interpretability of complex ML models. These barriers hinder the widespread adoption of ML in biological sensing applications. This review underscores the importance of interdisciplinary research and collaboration to address these challenges. Future work should focus on developing resource-efficient ML models, enhancing data transparency, and fostering innovations in areas such as data augmentation and model explainability to further advance the field. Key terms: Machine learning, biological sensing, biosensors, deep learning, environmental monitoring, medical diagnostics
- Research Article
4
- 10.5582/bst.2025.01013
- Apr 30, 2025
- Bioscience trends
This study investigates the use of machine learning (ML) models combined with a Synthetic Minority Over-sampling Technique (SMOTE) and its variants to predict perioperative pressure injuries (PIs) in an imbalanced dataset. PIs are a significant healthcare problem, often leading to prolonged hospitalization and increased medical costs. Conventional risk assessment scales are limited in their ability to predict PIs accurately, prompting the exploration of ML techniques to address this challenge.We utilized data from 7,292 patients admitted to a tertiary care hospital in Shanghai between May 2017 and July 2023, with a final dataset of 2,972 patients, including 158 with PIs. Seven ML algorithms-Support Vector Machine (SVM), Logistic Regression (LR), Random Forest (RF), Extreme Gradient Boosting (XGBoost), Extra Trees (ET), K-Nearest Neighbors (KNN), and Decision Trees (DT)-were used in conjunction with SMOTE, SMOTE+ENN, Borderline-SMOTE, ADASYN, and GAN to balance the dataset and improve model performance.Results revealed significant improvements in model performance when SMOTE and its variants were used. For instance, the XGBoost model hadan AUC of 0.996 with SMOTE, compared to 0.800 on raw data. SMOTE+ENN and Borderline-SMOTE further enhanced the models' ability to identify minority classes. External validation indicatedthat XGBoost, RF, and ET exhibited the highest stability and accuracy, with XGBoost having an AUC of 0.977. SHAP analysis revealed that factors such as anesthesia grade, age, and serum albumin levels significantly influenced model predictions.In conclusion, integrating SMOTE with ML algorithms effectively addressed a data imbalance and improved the prediction of perioperative PIs. Future work should focus on refining SMOTE techniques and exploring their application to larger, multi-center datasets to enhance the generalizability of these findings, and especially for diseaseswith a lowincidence.
- Research Article
95
- 10.1186/s12911-017-0566-6
- Dec 1, 2017
- BMC Medical Informatics and Decision Making
BackgroundPrior studies have demonstrated that cardiorespiratory fitness (CRF) is a strong marker of cardiovascular health. Machine learning (ML) can enhance the prediction of outcomes through classification techniques that classify the data into predetermined categories. The aim of this study is to present an evaluation and comparison of how machine learning techniques can be applied on medical records of cardiorespiratory fitness and how the various techniques differ in terms of capabilities of predicting medical outcomes (e.g. mortality).MethodsWe use data of 34,212 patients free of known coronary artery disease or heart failure who underwent clinician-referred exercise treadmill stress testing at Henry Ford Health Systems Between 1991 and 2009 and had a complete 10-year follow-up. Seven machine learning classification techniques were evaluated: Decision Tree (DT), Support Vector Machine (SVM), Artificial Neural Networks (ANN), Naïve Bayesian Classifier (BC), Bayesian Network (BN), K-Nearest Neighbor (KNN) and Random Forest (RF). In order to handle the imbalanced dataset used, the Synthetic Minority Over-Sampling Technique (SMOTE) is used.ResultsTwo set of experiments have been conducted with and without the SMOTE sampling technique. On average over different evaluation metrics, SVM Classifier has shown the lowest performance while other models like BN, BC and DT performed better. The RF classifier has shown the best performance (AUC = 0.97) among all models trained using the SMOTE sampling.ConclusionsThe results show that various ML techniques can significantly vary in terms of its performance for the different evaluation metrics. It is also not necessarily that the more complex the ML model, the more prediction accuracy can be achieved. The prediction performance of all models trained with SMOTE is much better than the performance of models trained without SMOTE. The study shows the potential of machine learning methods for predicting all-cause mortality using cardiorespiratory fitness data.
- Research Article
6
- 10.32628/ijsrset22924
- Mar 5, 2022
- International Journal of Scientific Research in Science, Engineering and Technology
Chronic Kidney Disease is one of the most critical illnesses nowadays and proper diagnosis is required as soon as possible. Machine learning technique has become reliable for medical treatment. With the help of a machine learning classifier algorithms, the doctor can detect the disease on time. For this perspective, Chronic Kidney Disease prediction has been discussed in this article. Chronic Kidney Disease dataset has been taken from the UCI repository. Seven classifier algorithms have been applied in this research such as artificial neural network, C5.0, Chi-square Automatic interaction detector, logistic regression, linear support vector machine with penalty L1 & with penalty L2 and random tree. The important feature selection technique was also applied to the dataset. For each classifier, the results have been computed based on the following factors given below: (i) full features, (ii) correlation-based feature selection, (iii) Wrapper method feature selection, (iv)Least absolute shrinkage and selection operator regression, (v) synthetic minority over- sampling technique with least absolute shrinkage and selection operator regression selected features, (vi) synthetic minority oversampling technique with full features. From the results, it is marked that LSVM with penalty L2 is giving the highest accuracy of 98.86% in synthetic minority over-sampling technique with full features. Along with accuracy, precision, recall, F- measure, area under the curve and GINI coefficient have been computed and compared results of various algorithms have been shown in the graph. Least absolute shrinkage and selection operator regression selected features with synthetic minority over-sampling technique gave the best after synthetic minority over-sampling technique with full features. In the synthetic minority over-sampling technique with least absolute shrinkage and selection operator selected features, again linear support vector machine gave the highest accuracy of 98.46%. Along with machine learning models one deep neural network has been applied on the same dataset and it has been noted that deep neural network achieved the highest accuracy of 99.6%.
- Research Article
1
- 10.14569/ijacsa.2025.01602132
- Jan 1, 2025
- International Journal of Advanced Computer Science and Applications
Cloud computing has transformed modern Information Technology (IT) infrastructures with its scalability and cost-effectiveness but introduces significant security risks. More-over, existing anomaly detection techniques are not well equipped to deal with the complexities of dynamic cloud environments. This systematic literature review shows the advancements in Machine Learning (ML) solutions for anomaly detection in cloud computing. The study categorizes ML approaches, examines the datasets and evaluation metrics utilized, and discusses their effectiveness and limitations. We analyze supervised, unsupervised, and hybrid ML models showing their advantages in dealing with a certain threat vector. It also discusses how advanced feature engineering, ensemble learning and real-time adaptability can improve detection accuracy and reduce false positives. Some key challenges, such as dataset diversity and computational efficiency, are highlighted, along with future research directions to improve ML based anomaly detection for robust and adaptive cloud security. Hybrid approaches are found to increase the accuracy reaching up to 99.85% and reduces the number of false positives. This review provides a comprehensive guide to researchers aiming to enhance anomaly detection in cloud environments.
- Research Article
4
- 10.5194/nhess-24-1913-2024
- Jun 6, 2024
- Natural Hazards and Earth System Sciences
Abstract. Landslides threaten human life and infrastructure, resulting in fatalities and economic losses. Monitoring stations provide valuable data for predicting soil movement, which is crucial in mitigating this threat. Accurately predicting soil movement from monitoring data is challenging due to its complexity and inherent class imbalance. This study proposes developing machine learning (ML) models with oversampling techniques to address the class imbalance issue and develop a robust soil movement prediction system. The dataset, comprising 2 years (2019–2021) of monitoring data from a landslide in Uttarakhand, has a 70:30 ratio of training and testing data. To tackle the class imbalance problem, various oversampling techniques, including the synthetic minority oversampling technique (SMOTE), K-means SMOTE, borderline-SMOTE, and adaptive SMOTE (ADASYN), were applied to the training dataset. Several ML models, namely random forest (RF), extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM), adaptive boosting (AdaBoost), category boosting (CatBoost), long short-term memory (LSTM), multilayer perceptron (MLP), and a dynamic ensemble, were trained and compared for soil movement prediction. A 5-fold cross-validation method was applied to optimize the ML models on the training data, and the models were tested on the testing set. Among these ML models, the dynamic ensemble model with K-means SMOTE performed the best in testing, with an accuracy, precision, and recall rate of 0.995, 0.995, and 0.995, respectively, and an F1 score of 0.995. Additionally, models without oversampling exhibited poor performance in training and testing, highlighting the importance of incorporating oversampling techniques to enhance predictive capabilities.
- Research Article
1
- 10.1136/bmjhci-2024-101189
- May 1, 2025
- BMJ Health & Care Informatics
ObjectiveThis study aimed to develop machine learning (ML) models to predict HIV status and assessed the factors associated with HIV infection among young men who have sex with men (MSM) under the Universal Health Coverage (UHC) programme in Thailand.MethodsYoung MSM aged 15–24 years who underwent HIV testing through the UHC programme from 2015 to 2022 were included. Data were divided into training (70%) and testing (30%) sets, with the Synthetic Minority Oversampling Technique (SMOTE) applied to address data set imbalance. ML models, including logistic regression, k-nearest neighbour (KNN), random forest, extreme gradient boosting (XGB) and AdaBoost, were used to predict HIV infection.ResultsAmong 146 813 young MSM, 11% were diagnosed with HIV. While KNN initially outperformed other ML models, the sensitivity of all models using the original data set was low due to imbalanced data. After applying SMOTE, the XGB model showed the best performance with an accuracy of 0.72, sensitivity of 0.73, specificity of 0.72 and the area under the curve of 0.72. The top predictors of HIV infection were the year of HIV testing (68%), age (55%) and targeted HIV testing (54%).DiscussionThis study demonstrates the potential of ML models, particularly XGB, in predicting HIV infection among young MSM in Thailand under the UHC programme. The application of SMOTE improved model sensitivity, addressing data imbalance and enhancing predictive accuracy.ConclusionsML models have the potential to enhance HIV risk assessment and inform targeted prevention strategies for high-risk populations.
- Research Article
31
- 10.2196/44081
- May 31, 2023
- Journal of Medical Internet Research
BackgroundLow birthweight (LBW) is a leading cause of neonatal mortality in the United States and a major causative factor of adverse health effects in newborns. Identifying high-risk patients early in prenatal care is crucial to preventing adverse outcomes. Previous studies have proposed various machine learning (ML) models for LBW prediction task, but they were limited by small and imbalanced data sets. Some authors attempted to address this through different data rebalancing methods. However, most of their reported performances did not reflect the models’ actual performance in real-life scenarios. To date, few studies have successfully benchmarked the performance of ML models in maternal health; thus, it is critical to establish benchmarks to advance ML use to subsequently improve birth outcomes.ObjectiveThis study aimed to establish several key benchmarking ML models to predict LBW and systematically apply different rebalancing optimization methods to a large-scale and extremely imbalanced all-payer hospital record data set that connects mother and baby data at a state level in the United States. We also performed feature importance analysis to identify the most contributing features in the LBW classification task, which can aid in targeted intervention.MethodsOur large data set consisted of 266,687 birth records across 6 years, and 8.63% (n=23,019) of records were labeled as LBW. To set up benchmarking ML models to predict LBW, we applied 7 classic ML models (ie, logistic regression, naive Bayes, random forest, extreme gradient boosting, adaptive boosting, multilayer perceptron, and sequential artificial neural network) while using 4 different data rebalancing methods: random undersampling, random oversampling, synthetic minority oversampling technique, and weight rebalancing. Owing to ethical considerations, in addition to ML evaluation metrics, we primarily used recall to evaluate model performance, indicating the number of correctly predicted LBW cases out of all actual LBW cases, as false negative health care outcomes could be fatal. We further analyzed feature importance to explore the degree to which each feature contributed to ML model prediction among our best-performing models.ResultsWe found that extreme gradient boosting achieved the highest recall score—0.70—using the weight rebalancing method. Our results showed that various data rebalancing methods improved the prediction performance of the LBW group substantially. From the feature importance analysis, maternal race, age, payment source, sum of predelivery emergency department and inpatient hospitalizations, predelivery disease profile, and different social vulnerability index components were important risk factors associated with LBW.ConclusionsOur findings establish useful ML benchmarks to improve birth outcomes in the maternal health domain. They are informative to identify the minority class (ie, LBW) based on an extremely imbalanced data set, which may guide the development of personalized LBW early prevention, clinical interventions, and statewide maternal and infant health policy changes.
- Dissertation
- 10.32657/10356/173687
- Jan 1, 2024
In recent years, the implementation of digital twin (DT) as a digital replica of the physical asset has matured significantly in smart manufacturing with the advancement of digital technologies. At the same time, for water treatment facilities which are critical infrastructures, DT is still in its infancy for real-world applications. Therefore, there is a pressing need for research that can accelerate the DT development for these critical infrastructures. The aim of this study is to improve the performance of DT for water treatment facilities using probabilistic assessment. The scope of research focuses on two main directions: first, utilizing probabilistic machine learning (ML) models to assist in real-time anomaly detection of DT, and second, developing data assimilation methods for probabilistic ML models to obtain optimized system states for process control. Anomaly detection is crucial for water treatment facilities as they can be susceptible to cyber-physical attacks that negatively impact the system's functionality, while data assimilation can improve the robustness of real-time monitoring in noisy environments by assimilating ML predictions with observations. A combined anomaly detection framework (CADF) was first developed for DT applications using probabilistic ML models. A prototype water treatment testbed facility was utilized to verify the anomaly detection framework by simulating various types of security attacks. CADF utilizes a Programmable Logic Controller (PLC)-based whitelist system to detect anomalies targeting the actuators, and a probabilistic ML model with corresponding assessment to detect anomalies targeting the sensors in this fully operational, scaled-down testbed facility. The results showed that CADF could successfully detect various attacks on the DT and reduce false alarms significantly when compared with other methods. To further enhance the performance of CADF, a real-time data processing framework was developed to update ML models in the DT by selecting suitable update intervals and training datasets to maximize their overall accuracy. It was shown to synchronize model updates and predictions effectively, leading to a significant reduction in errors. A data assimilation approach named Probabilistic Optimal Interpolation (POI) was also developed to combine the predictions from probabilistic ML models and real-time observations. The quantification of the respective uncertainties is directly included within the probabilistic ML model itself. As an application example, the performance of POI was tested using a multi-scale Lorenz 96 chaos system in both stationary and nonstationary environments. The POI implementation was able to reduce uncertainty in both environments and serve as a compromise in scenarios where the noise level of the environment was unclear. The performance of POI under scenarios with missing values was also evaluated by masking the test datasets with different missingness rates. The impact from random missing values was found negligible and assimilation was still suggested at missing points. In summary, probabilistic approaches and frameworks were developed in this study for DT of water treatment facilities, with a particular emphasis on anomaly detection and data assimilation. They were shown to be effective in enhancing the DT performance and potentially leading to more robust and secure systems in the future.