Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Integrating Multi-Variable Driving Factors to Improve Land Use & Land Cover Classification Accuracy using Machine Learning Approaches: A Case Study from Lombok Island

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Accurate classification of land cover is essential for effective land management and environmental monitoring. This study aimed to enhance land cover classification for Lombok Island using advanced machine learning algorithms. The models employed include Random Forest, Gradient Boosting, Decision Tree, and Naive Bayes, integrating a wide range of variables, such as Landsat satellite imagery, spectral indices, physiographic, climatic, and socio-economic data. Among these, Random Forest demonstrated the highest model accuracy at 82%, followed by Gradient Boosting at 80%, Decision Tree at 73%, and Naïve Bayes at 61%. In field validation assessments, comparing the predictions of these machine learning models with ground truth data, Random Forest was the most reliable, achieving an overall accuracy of 88%. This superior performance is largely due to the multi-variable approach, which allows the model to mitigate issues like cloud cover in satellite images. The key variables that significantly influenced the land cover classification on Lombok Island include proximity to settlements, temperature, and distance to roads. These results provide essential insights for land management strategies, enabling policymakers and stakeholders to make informed decisions on sustainable development, urban planning, and environmental conservation in rapidly changing landscapes.

Similar Papers
  • PDF Download Icon
  • Research Article
  • Cite Count Icon 165
  • 10.3390/rs11020164
Classification of Land Cover, Forest, and Tree Species Classes with ZiYuan-3 Multispectral and Stereo Data
  • Jan 16, 2019
  • Remote Sensing
  • Zhuli Xie + 4 more

The global availability of high spatial resolution images makes mapping tree species distribution possible for better management of forest resources. Previous research mainly focused on mapping single tree species, but information about the spatial distribution of all kinds of trees, especially plantations, is often required. This research aims to identify suitable variables and algorithms for classifying land cover, forest, and tree species. Bi-temporal ZiYuan-3 multispectral and stereo images were used. Spectral responses and textures from multispectral imagery, canopy height features from bi-temporal stereo imagery, and slope and elevation from the stereo-derived digital surface model data were examined through comparative analysis of six classification algorithms including maximum likelihood classifier (MLC), k-nearest neighbor (kNN), decision tree (DT), random forest (RF), artificial neural network (ANN), and support vector machine (SVM). The results showed that use of multiple source data—spectral bands, vegetation indices, textures, and topographic factors—considerably improved land-cover and forest classification accuracies compared to spectral bands alone, which the highest overall accuracy of 84.5% for land cover classes was from the SVM, and, of 89.2% for forest classes, was from the MLC. The combination of leaf-on and leaf-off seasonal images further improved classification accuracies by 7.8% to 15.0% for land cover classes and by 6.0% to 11.8% for forest classes compared to single season spectral image. The combination of multiple source data also improved land cover classification by 3.7% to 15.5% and forest classification by 1.0% to 12.7% compared to the spectral image alone. MLC provided better land-cover and forest classification accuracies than machine learning algorithms when spectral data alone were used. However, some machine learning approaches such as RF and SVM provided better performance than MLC when multiple data sources were used. Further addition of canopy height features into multiple source data had no or limited effects in improving land-cover or forest classification, but improved classification accuracies of some tree species such as birch and Mongolia scotch pine. Considering tree species classification, Chinese pine, Mongolia scotch pine, red pine, aspen and elm, and other broadleaf trees as having classification accuracies of over 92%, and larch and birch have relatively low accuracies of 87.3% and 84.5%. However, these high classification accuracies are from different data sources and classification algorithms, and no one classification algorithm provided the best accuracy for all tree species classes. This research implies the same data source and the classification algorithm cannot provide the best classification results for different land cover classes. It is necessary to develop a comprehensive classification procedure using an expert-based approach or hierarchical-based classification approach that can employ specific data variables and algorithm for each tree species class.

  • Research Article
  • 10.4236/ojapps.2026.161005
An Explainable Wavelet-Based Feature Decomposition and Machine Learning Framework for Land Cover Classification
  • Jan 1, 2026
  • Open Journal of Applied Sciences
  • Saviour Mantey + 2 more

Accurate land cover classification is essential for environmental monitoring, urban planning, and resource management. Conventional classifiers trained on raw spectral bands are often limited by noise, inter-class spectral similarity, and intra-class variability. This study introduces a wavelet-based feature decomposition and machine learning framework to address these challenges. Landsat-8 Operational Land Imager (OLI) Level-2 surface reflectance data were pre-processed and decomposed using a one-dimensional discrete wavelet transform to isolate low- and high-frequency components. The decomposed features were concatenated with raw bands to form an enriched dataset, which was used to train and validate three supervised classifiers: Random Forest (RF), Gradient Boosting (GBM), and Decision Tree (DT). Model performance was evaluated using 5-fold cross-validation, and the best-performing model, the Random Forest Wavelet Transform (RF_WT), achieved the highest macro-F1 score (0.865). Explainable AI (Gini importance, permutation importance, SHAP (SHapley Additive exPlanations) values) confirmed the complementary role of raw and decomposed features, with wavelet-derived detail components in the Short-Wave Infrared (SWIR) and visible bands strongly influencing classification. The RF_WT model, when compared to RF, DT and GBM classifiers trained on raw spectral bands across all classes, achieved the highest macro-averaged accuracy values (UA = 0.876, PA = 0.864, F1 = 0.865), compared with RF (UA = 0.864, PA = 0.852, F1 = 0.853), GBM (UA = 0.849, PA = 0.832, F1 = 0.832), and DT (UA = 0.815, PA = 0.808, F1 = 0.809). The approach is valuable in support of ecological monitoring and resource management. However, challenges in settlement classification highlight the limitations of medium-resolution data. Future work could integrate higher-resolution imagery or multi-temporal data for improved performance.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 88
  • 10.3389/fcvm.2022.839379
Machine Learning Approaches for Predicting Hypertension and Its Associated Factors Using Population-Level Data From Three South Asian Countries
  • Mar 31, 2022
  • Frontiers in Cardiovascular Medicine
  • Sheikh Mohammed Shariful Islam + 11 more

BackgroundHypertension is the most common modifiable risk factor for cardiovascular diseases in South Asia. Machine learning (ML) models have been shown to outperform clinical risk predictions compared to statistical methods, but studies using ML to predict hypertension at the population level are lacking. This study used ML approaches in a dataset of three South Asian countries to predict hypertension and its associated factors and compared the model's performances.MethodsWe conducted a retrospective study using ML analyses to detect hypertension using population-based surveys. We created a single dataset by harmonizing individual-level data from the most recent nationally representative Demographic and Health Survey in Bangladesh, Nepal, and India. The variables included blood pressure (BP), sociodemographic and economic factors, height, weight, hemoglobin, and random blood glucose. Hypertension was defined based on JNC-7 criteria. We applied six common ML-based classifiers: decision tree (DT), random forest (RF), gradient boosting machine (GBM), extreme gradient boosting (XGBoost), logistic regression (LR), and linear discriminant analysis (LDA) to predict hypertension and its risk factors.ResultsOf the 8,18,603 participants, 82,748 (10.11%) had hypertension. ML models showed that significant factors for hypertension were age and BMI. Ever measured BP, education, taking medicine to lower BP, and doctor's perception of high BP was also significant but comparatively lower than age and BMI. XGBoost, GBM, LR, and LDA showed the highest accuracy score of 90%, RF and DT achieved 89 and 83%, respectively, to predict hypertension. DT achieved the precision value of 91%, and the rest performed with 90%. XGBoost, GBM, LR, and LDA achieved a recall value of 100%, RF scored 99%, and DT scored 90%. In F1-score, XGBoost, GBM, LR, and LDA scored 95%, while RF scored 94%, and DT scored 90%. All the algorithms performed with good and small log loss values <6%.ConclusionML models performed well to predict hypertension and its associated factors in South Asians. When employed on an open-source platform, these models are scalable to millions of people and might help individuals self-screen for hypertension at an early stage. Future studies incorporating biochemical markers are needed to improve the ML algorithms and evaluate them in real life.

  • Research Article
  • 10.62222/bvch4721
Profitability and SDG 13-Climate Action: The Moderating Role of Environmental Management Training: A Machine Learning Approach
  • Jun 30, 2025
  • Journal of Business Sectors
  • Hassan Raza + 2 more

Research background: This study examined the critical intersection between corporate profitability and adherence to Sustainable Development Goal (SDG) 13 - Climate Action, highlighting the importance of environmental management training within organizations. As the global business landscape evolves, it is essential to understand how internal sustainability practices influence financial outcomes. This research addressed the gap in understanding the dynamic interplay between corporate financial performance and strategic commitment to environmental sustainability. Purpose of the article: The primary objective of this research was to analyze the moderating role of environmental management training on the relationship between firm profitability and support for SDG 13. It attempted to elucidate how structured training initiatives can improve corporate financial performance while promoting robust climate action strategies. Methods: This study used a comprehensive panel dataset of 31,346 firms in 67 countries from 2013 to 2022. Advanced machine learning algorithms, including decision tree, random forest, gradient boosting, and logistic regression, were employed to analyze the impact of environmental management training on corporate profitability and sustainability initiatives. These methods allow for a nuanced understanding of the complex, non-linear interactions among the variables studied, providing deep insights into the dynamics at play. Findings &amp; Value added: The results indicate that certain predictive models, such as Decision Tree and Random Forest, initially faced challenges like overfitting. However, incorporating environmental management training variables improved their robustness, highlighting the importance of variable selection in sustainability analytics. Among the models tested, Gradient Boosting demonstrated a strong balance between precision and recall, making it particularly effective for predicting corporate engagement in climate initiatives. The incorporation of machine learning provides a novel methodological perspective that deepens our understanding of how profitability metrics can influence and enhance corporate sustainability efforts. This research adds value to the discourse on sustainable business practices by providing robust empirical evidence and methodological innovations that can guide policymakers and business leaders in crafting strategies that promote sustainable development. Moreover, this study aligns with the journal’s focus by offering innovative approaches to economic policy and enhancing understanding of the intersections between corporate strategy and sustainable business practices.

  • Research Article
  • Cite Count Icon 1
  • 10.52783/jisem.v10i19s.3009
Quantum Computing Base Cybersecurity Mathematical Model Development for Geographically Underdeveloped Areas using Multiple Zonal Approaches using AIML Techniques for Stoppage of Different Types of Attacks
  • Mar 12, 2025
  • Journal of Information Systems Engineering and Management
  • Nandini G.S

Our methodology utilizes a supervised learning approach, employing Random Forest and Gradient Boosting Machines (GBM) trained on a comprehensive dataset that includes email headers, content, and sender behavior. This approach allows our models to discern complex patterns associated with phishing attempts, achieving a 92% detection rate, a substantial improvement over the traditional signature-based methods' 65% rate. Additionally, we integrated NLP techniques, specifically Word2Vec and GloVe, to extract semantic features from email content, enhancing our system's ability to identify malicious intent. The incorporation of NLP not only improves the precision of phishing detection by an additional 15% compared to conventional methods but also emphasizes the importance of semantic analysis in cybersecurity. This enhancement is crucial for understanding the subtle cues within email content that may indicate phishing, offering a more robust and effective defense mechanism for rural areas. By combining supervised learning with quantum computing and NLP, our approach addresses the significant gaps in traditional cybersecurity methods. This multi-layered strategy ensures a more reliable and efficient way to safeguard rural communities from the increasing threat of cyber attacks. The advanced AI techniques employed here leverage both the predictive power of machine learning and the nuanced understanding of language provided by NLP, setting a new standard in cybersecurity practices. The results of our study highlight the effectiveness of the proposed methodology, demonstrating a potential to markedly improve cybersecurity in resource-constrained rural environments. With a 92% phishing detection rate and an increase in precision through the use of NLP, our approach promises a significant advancement in the protection against cyber threats for rural areas, offering a comprehensive and scalable solution. This research presents an innovative multi-layered AI approach, utilizing quantum computing to enhance cybersecurity in rural areas vulnerable to phishing threats. The paper details the integration of sophisticated machine learning techniques—Random Forest and Gradient Boosting Machines (GBM)—with Natural Language Processing (NLP) tools like Word2Vec and GloVe, achieving significant improvements in phishing detection rates. Through a comprehensive analysis of existing cybersecurity strategies and the limitations of traditional signature-based detection methods, this study proposes a robust solution tailored for rural settings such as Siddlagatta, Chikkaballapur, and Devanahalli. By incorporating quantum computing, the approach not only overcomes the constraints of classical computing but also leverages the predictive prowess of AI to offer a more reliable and effective defense against cyber threats. The results demonstrate a promising increase in detection rates, underscoring the potential of this quantum-enhanced, AI-driven strategy to significantly bolster cybersecurity in resource-limited rural environments. Introduction : Cybersecurity in rural areas remains a pivotal concern, exacerbated by limited access to sophisticated technological resources and infrastructure. This paper introduces an advanced multi-layered artificial intelligence (AI) approach, utilizing quantum computing to enhance phishing threat detection in rural environments. Focusing on regions like Siddlagatta, Chikkaballapur, and Devanahalli, the study integrates supervised learning algorithms—Random Forest and Gradient Boosting Machines (GBM)—with Natural Language Processing (NLP) techniques to improve the detection and analysis of phishing attempts. By leveraging machine learning to surpass traditional signature-based methods, this approach significantly boosts detection rates, presenting a tailored, effective solution to protect these vulnerable communities against evolving cyber threats.. Objectives : The objectives of this research are to develop and implement a multi-layered artificial intelligence (AI) approach, utilizing quantum computing to enhance the detection of phishing threats in rural areas. Specifically, the study aims to address the limitations of traditional signature-based detection methods by integrating advanced machine learning algorithms such as Random Forest and Gradient Boosting Machines (GBM) with Natural Language Processing (NLP) techniques. This integration seeks to improve the precision of identifying malicious intent in email communications by analyzing semantic features. The research also explores the effectiveness of these AI techniques in rural settings where cybersecurity resources are scarce, aiming to provide a more robust and efficient solution that can significantly reduce the incidence of phishing attacks in these vulnerable communities. Methods : The proposed methodology entails the development of a web-based platform that melds social networking functionalities with sophisticated agricultural tools and services. By utilizing user profiles, the system effectively categorizes key stakeholders such as farmers, suppliers, experts, and policymakers to foster focused engagement and collaborative efforts. The integration of data from IoT sensors, satellite imagery, and user contributions is channeled into a central system that supports real-time analysis and informed decision-making. Moreover, the platform employs algorithms designed to align stakeholders with pertinent resources, market possibilities, and professional advice. Enhanced communication features like forums, direct messaging, and video conferencing are incorporated to promote interactive exchanges among users. A pilot phase involving select agricultural communities will be initiated to evaluate the practicality and impact of the framework, with subsequent adjustments driven by user feedback and analytic assessments. The ultimate goal of this framework is to boost connectivity, facilitate the efficient distribution of resources, and empower all involved parties through a scalable and intuitive interface. This approach not only aims to revolutionize the way agricultural communities interact and operate but also seeks to provide a robust foundation for continuous growth and innovation in the sector. Results : The simulated results of the study demonstrate a significant enhancement in phishing detection capabilities through the integration of a multi-layered AI approach in rural settings. The deployment of advanced machine learning algorithms, such as Random Forest and Gradient Boosting Machines (GBM), along with Natural Language Processing (NLP) techniques, notably increased the phishing detection rate to 92%, a substantial improvement over the 65% detection rate achieved by traditional signature-based methods. Additionally, the incorporation of NLP through tools like Word2Vec and GloVe improved the precision of identifying malicious intent by an additional 15%, emphasizing the effectiveness of semantic analysis in distinguishing phishing attempts. These results highlight the potential of combining machine learning and quantum computing to address the unique cybersecurity challenges faced in rural areas, providing a robust solution that significantly enhances the detection and prevention of phishing threats.. Conclusions : The research presented in this paper successfully demonstrates the efficacy of a multi-layered AI approach in significantly enhancing cybersecurity against phishing threats in rural areas. By integrating advanced machine learning algorithms with Natural Language Processing techniques and quantum computing, the study achieved a notable increase in phishing detection rates, outperforming traditional signature-based methods with a detection rate of 92%. This approach not only addresses the limitations inherent in existing cybersecurity measures but also tailors its strategy to the unique challenges posed by the limited resources and infrastructure in rural environments. The integration of semantic analysis through NLP further enhanced the precision of threat detection, providing a more nuanced understanding of malicious intent. Overall, the study underscores the potential of sophisticated AI technologies to transform cybersecurity practices in underserved areas, ensuring more effective protection against evolving cyber threats.

  • Research Article
  • Cite Count Icon 100
  • 10.1007/s11657-020-00802-8
Application of machine learning approaches for osteoporosis risk prediction in postmenopausal women.
  • Oct 23, 2020
  • Archives of Osteoporosis
  • Jae-Geum Shim + 6 more

Osteoporosis is a silent disease until it results in fragility fractures. However, early diagnosis of osteoporosis provides an opportunity to detect and prevent fractures. We aimed to develop machine learning approaches to achieve high predictive ability for osteoporosis risk that could help primary care providers identify which women are at increased risk of osteoporosis and should therefore undergo further testing with bone densitometry. We included all postmenopausal Korean women from the Korea National Health and Nutrition Examination Surveys (KNHANES V-1, V-2) conducted in 2010 and 2011. Machine learning models using methods such as the k-nearest neighbors (KNN), decision tree (DT), random forest (RF), gradient boosting machine (GBM), support vector machine (SVM), artificial neural networks (ANN), and logistic regression (LR) were developed to predict osteoporosis risk. We analyzed the effect of applying the machine learning algorithms to the raw data and featuring the selected data only where the statistically significant variables were included as model inputs. The accuracy, sensitivity, specificity, and area under the receiver operating characteristic curve (AUROC) were used to evaluate performance among the seven models. A total of 1792 patients were included in this study, of which 613 had osteoporosis. The raw data consisted of 19 variables and achieved performances (in terms of AUROCs) of 0.712, 0.684, 0.727, 0.652, 0.724, 0.741, and 0.726 for KNN, DT, RF, GBM, SVM, ANN, and LR with fivefold cross-validation, respectively. The feature selected data consisted of nine variables and achieved performances (in terms of AUROCs) of 0.713, 0.685, 0.734, 0.728, 0.728, 0.743, and 0.727 for KNN, DT, RF, GBM, SVM, ANN, and LR with fivefold cross-validation, respectively. In this study, we developed and compared seven machine learning models to accurately predict osteoporosis risk. The ANN model performed best when compared to the other models, having the highest AUROC value. Applying the ANN model in the clinical environment could help primary care providers stratify osteoporosis patients and improve the prevention, detection, and early treatment of osteoporosis.

  • Research Article
  • Cite Count Icon 3
  • 10.22495/rgcv15i1p3
Traditional or advanced machine learning approaches: Which one is better for housing price prediction and uncertainty risk reduction?
  • Jan 1, 2025
  • Risk Governance and Control: Financial Markets and Institutions
  • Long Phi Tran + 3 more

Predicting housing prices is particularly of interest to many scholars and policymakers. However, housing prices are highly volatile and difficult to predict. This study used both traditional and advanced machine learning (ML) approaches to address the issue of housing price prediction. This study involves and compares the predictive power between advanced ML models, including random forest, gradient boosting, k-nearest neighbors (KNN), bagged classification and regression trees (CART), and traditional ML models based on linear regression and its modifications. Notably, in this study, we employed both performance metrics, including the mean absolute error (MAE), root mean square error (RMSE), coefficient of determination (R2), and k-fold cross-validation (CV) procedure in order to investigate the predictive performance of each model. Empirically, based on a dataset comprising 78,704 real estate sales in Hanoi, Vietnam, we find that advanced ML approaches outperform traditional approaches. Specifically, advanced ML models enhance the accuracy of house price prediction and the decision-making process related to housing buying and selling activities. Our findings also reveal that among advanced ML algorithms, the random forest algorithm performs better than the other models in predicting housing prices.

  • Conference Article
  • Cite Count Icon 4
  • 10.2523/iptc-24998-ms
Data-Driven Prediction of Storage Column Height for H2-Brine Systems: Accelerating Underground Hydrogen Storage
  • Feb 17, 2025
  • Aneeq Nasir Janjua + 4 more

A practical solution to energy transition and the increasing demand for energy is underground hydrogen storage (UHS). The contribution of hydrogen (H2) as a clean energy source has proven to be an effective substitute for future use to meet the net-zero target and reduce anthropogenic greenhouse gas emissions. One of the most important factors affecting H2 displacement and storage capacity under geological circumstances is storage column height. The objective of this study is to underscore the importance of large-scale H2 storage and use reliable machine learning algorithms to evaluate and predict the H2 storage column height under varied thermophysical and salinity conditions. In this study, the dataset of 540 datapoints for the evaluation and prediction of storage column height is generated, which involves three main parameters: density difference (Δρ), interfacial tension (IFT) and contact angle (θ). The correlation of contact angles against various reservoir depths is used and H2 storage column height is evaluated. Thermophysical conditions include pressures (0.1-20 MPa), temperatures (25-70°C), and salinities including deionized water, seawater and brines of 1 and 3 molar concentrations for various salts (NaCl, KCl, MgCl2, CaCl2, and Na2SO4) from our experimental data. The H2 storage column height (h) is predicted using three machine learning (ML) models, viz., random forest (RF), decision tree (DT) and gradient boosting (GB). Statistical data analysis is performed to generate the distribution of dataset and correlation coefficient is calculated while feature importance is determined to identify the relationship of each input parameter with output parameter using Pearson, Spearman, and Kendall models. RF and GB, as demonstrated in this study, have shown promising results in providing accurate predictions while maintaining generalizability. Various error assessment metrics including MSE, RMSE, MAPE and R2 are utilized for the evaluation. Prediction of column height resulted in R2 values of 0.995 for training and 0.999 for testing with RF model. Whereas the GB model also resulted in superior performance with R2 values of 0.997 during the training phase and 0.995 during the testing phase. However, the DT model resulted in R2 values of 1 and 0.994 during the training and testing phases respectively. While MSE value of 0 is obtained for DT model which indicated overfitting. The findings of this study suggest that data-driven ML models can be a powerful tool for accurately predicting the H2 storage column height and can be effectively used to determine the displacement of H2 and storage capacity, reducing the time and cost associated with determination using traditional methods. In addition, advanced ML algorithms can be explored in the future to overcome the challenges pertinent to the determination of storage column height.

  • Research Article
  • 10.56294/dm2025755
Classifying Dental Care Providers Through Machine Learning with Features Ranking
  • Apr 7, 2025
  • Data and Metadata
  • Mohammad Subhi Al-Batah Al-Batah + 4 more

This study investigates the application of machine learning (ML) models for classifying dental providers into two categories—standard rendering providers and safety net clinic (SNC) providers—using a 2018 dataset of 24,300 instances with 20 features. The dataset, characterized by high missing values (38.1%), includes service counts (preventive, treatment, exams), delivery systems (FFS, managed care), and beneficiary demographics. Feature ranking methods such as information gain, Gini index, and ANOVA were employed to identify critical predictors, revealing treatment-related metrics (TXMT_USER_CNT, TXMT_SVC_CNT) as top-ranked features. Twelve ML models, including k-Nearest Neighbors (kNN), Decision Trees, Support Vector Machines (SVM), Stochastic Gradient Descent (SGD), Random Forest, Neural Networks, and Gradient Boosting, were evaluated using 10-fold cross-validation. Classification accuracy was tested across incremental feature subsets derived from rankings. The Neural Network achieved the highest accuracy (94.1%) using all 20 features, followed by Gradient Boosting (93.2%) and Random Forest (93.0%). Models showed improved performance as more features were incorporated, with SGD and ensemble methods demonstrating robustness to missing data. Feature ranking highlighted the dominance of treatment service counts and annotation codes in distinguishing provider types, while demographic variables (AGE_GROUP, CALENDAR_YEAR) had minimal impact. The study underscores the importance of feature selection in enhancing model efficiency and accuracy, particularly in imbalanced healthcare datasets. These findings advocate for integrating feature-ranking techniques with advanced ML algorithms to optimize dental provider classification, enabling targeted resource allocation for underserved populations.

  • Research Article
  • 10.22317/jcms.v12i1.2093
Prevalence and Associated Factors of Hypertension Among Adolescents: A Machine Learning Approach
  • Feb 26, 2026
  • Journal of Contemporary Medical Sciences
  • Mohammed Saad Abdullah + 1 more

Objective: This study aims to estimate the prevalence of HTN among high school students in Erbil City, Kurdistan Region, Iraq, and to identify associated factors using a machine learning approach. Methods: A school-based cross-sectional study was conducted (n = 1619). Eight distinct supervised learning algorithms included Logistic Regression (LR), k-Nearest Neighbors (kNN), Decision Tree (DT), Random Forest (RF), Gradient Boosting Machine (GBM), Extreme Gradient Boosting (XGBoost), Naïve Bayes (NB), and Support Vector Machine (SVM), were implemented to compare the predictive performance. The performance of each model was assessed using the following metrics: Area Under the Receiver Operating Characteristic Curve (AUC-ROC), Area Under the Precision-Recall Curve (AUC-PR), Balanced Accuracy, Precision, Recall, Specificity, and F1 score. The best-performing model was interpreted using SHapley Additive exPlanations. Results: The prevalence of HTN was 11.3% (95% CI: 9.8% - 13.0%). The highest (AUC-ROC) value (0.8306) was observed for the LR model. However, the XGBoost, RF, and GBM models exhibited slightly lower AUC-ROC values (0.8173, 0.8078, and 0.8041, respectively). Key factors for HTN prediction included higher salt intake, higher BMI, older age, higher sedentary behavior (hr/day), lower physical activity (days/week), positive family history of HTN, female sex, lower vegetable intake, lower sleep duration (hr/night), and lower physical activity (min/day). Conclusion: This study demonstrated that machine learning algorithms, particularly LR and XGBoost, provide high discriminative power for predicting adolescent hypertension. Future research should focus on validating these models across diverse geographic cohorts to ensure generalizability.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 37
  • 10.3390/ijgi9050329
Decision Tree Algorithms for Developing Rulesets for Object-Based Land Cover Classification
  • May 19, 2020
  • ISPRS International Journal of Geo-Information
  • Darius Phiri + 4 more

Decision tree (DT) algorithms are important non-parametric tools used for land cover classification. While different DTs have been applied to Landsat land cover classification, their individual classification accuracies and performance have not been compared, especially on their effectiveness to produce accurate thresholds for developing rulesets for object-based land cover classification. Here, the focus was on comparing the performance of five DT algorithms: Tree, C5.0, Rpart, Ipred, and Party. These DT algorithms were used to classify ten land cover classes using Landsat 8 images on the Copperbelt Province of Zambia. Classification was done using object-based image analysis (OBIA) through the development of rulesets with thresholds defined by the DTs. The performance of the DT algorithms was assessed based on: (1) DT accuracy through cross-validation; (2) land cover classification accuracy of thematic maps; and (3) other structure properties such as the sizes of the tree diagrams and variable selection abilities. The results indicate that only the rulesets developed from DT algorithms with simple structures and a minimum number of variables produced high land cover classification accuracies (overall accuracy &gt; 88%). Thus, algorithms such as Tree and Rpart produced higher classification results as compared to C5.0 and Party DT algorithms, which involve many variables in classification. This high accuracy has been attributed to the ability to minimize overfitting and the capacity to handle noise in the data during training by the Tree and Rpart DTs. The study produced new insights on the formal selection of DT algorithms for OBIA ruleset development. Therefore, the Tree and Rpart algorithms could be used for developing rulesets because they produce high land cover classification accuracies and have simple structures. As an avenue of future studies, the performance of DT algorithms can be compared with contemporary machine-learning classifiers (e.g., Random Forest and Support Vector Machine).

  • Research Article
  • Cite Count Icon 6
  • 10.32629/jai.v6i2.623
Experiences of sexual minorities on social media: A study of sentiment analysis and machine learning approaches
  • Aug 4, 2023
  • Journal of Autonomous Intelligence
  • Peter Appiahene + 3 more

&lt;p&gt;Nowadays, social media has become a forum for people to express their views on issues such as sexual orientation, legislation, and taxes. Sexual orientation refers to individuals with whom you are attracted and wish to be engaged. In the world, many people are regarded as having different sexual orientations. People categorized as lesbian, gay, bisexual, transgender, queer, and many more (LGBTQ+) have many sexual orientations. Because of the public stigmatization of LGBTQ+ persons, many turn to social media to express themselves, sometimes anonymously. The present study aims to use natural language processing (NLP) and machine learning (ML) approaches to assess the experiences of LGBTQ+ persons. To train the data, the study used lexicon-based sentiment analysis (SA) and six distinct machine classifiers, including logistic regression (LR), support vector machine (SVM), naïve bayes (NB), decision tree (DT), random forest (RF), and gradient boosting (GB). Individuals are positive about LGBTQ concerns, according to the SA results; yet, prejudice and harsh statements against the LGBTQ people persist in many regions where they live, according to the negative sentiment ratings. Furthermore, using LR, SVM, NB, DT, RF, and GB, the ML classifiers attained considerable accuracy values of 97%, 96%, 88%, 100%, 92%, and 91%, respectively. The performance assessment metrics used obtained significant recall and precision values. This study will assist the government, non-governmental organizations, and rights advocacy groups make educated decisions about LGBTQ+ concerns in order to ensure a sustainable future and peaceful coexistence.&lt;/p&gt;

  • Research Article
  • Cite Count Icon 130
  • 10.1016/j.mineng.2018.04.010
An intelligent modelling framework for mechanical properties of cemented paste backfill
  • Apr 24, 2018
  • Minerals Engineering
  • Chongchong Qi + 3 more

An intelligent modelling framework for mechanical properties of cemented paste backfill

  • Research Article
  • Cite Count Icon 3
  • 10.56536/jbahs.v5i1.100
Application of Machine Learning for Optimizing Oil Well Production and Reservoir Management: A Simulation-Based Approach
  • Feb 13, 2025
  • Journal of Biological and Allied Health Sciences
  • Mohsin Saleem

Background: The oil and gas industry requires efficient reservoir management and accurate production forecasting to optimize operations and reduce costs. Traditional physics-based models, though reliable, are computationally intensive and require domain expertise. Machine learning (ML) offers a data-driven approach to predict production trends, optimize operational strategies, and enhance decision-making. This study evaluates various ML models, including regression, decision trees, gradient boosting machines (GBM), and deep learning, to determine their effectiveness in oil well production forecasting. Methods: A synthetic dataset simulating reservoir conditions, production histories, and operational parameters was used to train ML models. Linear regression, decision trees, random forests, GBM, and deep learning models were tested. Performance was measured using Root Mean Squared Error (RMSE) and Mean Absolute Error (MAE). Hyperparameter tuning and cross-validation were applied to improve model accuracy, and feature importance analysis was conducted to identify key factors influencing production. Results: GBM achieved the highest accuracy, with an RMSE of 3.5% and an MAE of 2.1%, outperforming other models in production forecasting. Deep learning models captured complex patterns but required high computational resources. Random forests showed strong generalization, making them effective for noisy datasets, while linear regression struggled with non-linearity. Overall, ML models improved forecasting accuracy and enabled real-time optimization of reservoir operations. Conclusion: ML models significantly enhance oil well production forecasting and reservoir management. GBM proved to be the most effective, balancing accuracy and efficiency. Integrating ML into oil well operations can reduce costs and improve decision-making. Future research should focus on real-world datasets and hybrid ML approaches to further refine predictive capabilities.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 21
  • 10.3390/bioengineering10030277
Screening for Osteoporosis from Blood Test Data in Elderly Women Using a Machine Learning Approach
  • Feb 21, 2023
  • Bioengineering
  • Atsuyuki Inui + 9 more

The diagnosis of osteoporosis is made by measuring bone mineral density (BMD) using dual-energy X-ray absorptiometry (DXA). Machine learning, one of the artificial intelligence methods, was used to predict low BMD without using DXA in elderly women. Medical records from 2541 females who visited the osteoporosis clinic were used in this study. As hyperparameters for machine learning, patient age, body mass index (BMI), and blood test data were used. As machine learning models, logistic regression, decision tree, random forest, gradient boosting trees, and lightGBM were used. Each model was trained to classify and predict low-BMD patients. The model performance was compared using a confusion matrix. The accuracy of each trained model was 0.772 in logistic regression, 0.739 in the decision tree, 0.775 in the random forest, 0.800 in gradient boosting, and 0.834 in lightGBM. The area under the curve (AUC) was 0.595 in the decision tree, 0.673 in logistic regression, 0.699 in the random forest, 0.840 in gradient boosting, and 0.961, which was the highest, in the lightGBM model. Important features were BMI, age, and the number of platelets. Shapley additive explanation scores in the lightGBM model showed that BMI, age, and ALT were ranked as important features. Among several machine learning models, the lightGBM model showed the best performance in the present research.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant