Domestic violence in Nepal: Insights from machine learning-based prediction.
Conducting surveys on domestic violence across diverse countries, particularly in lower-middle-income nations like Nepal, poses significant challenges in understanding and addressing the multifaceted dynamics involved in domestic violence research. However, integrating machine learning can help uncover patterns and predictive factors. Therefore, this study aimed to evaluate and compare machine-learning models to identify population-level risk patterns of domestic violence associated with male demographic characteristics using nationally representative data from Nepal. We utilized nationally representative data from the Nepal Demographic and Health Surveys (DHS) conducted in 2016 and 2022. A total of 7,813 observations were analyzed. The outcome variable captured whether women reported experiencing any form of physical or sexual violence. Data preprocessing and analysis were conducted using Stata and Python, with machine learning models implemented through the PyCaret framework. Multiple algorithms were evaluated based on performance metrics including accuracy, precision, recall, F1-score, and AUC. Significant demographic shifts were observed between 2016 and 2022, including an increase in husbands with only primary education (from 23.2% to 42.52%) and rising rates of alcohol consumption. Among all models tested, LDA achieved the highest accuracy (74.61%) and F1-score (0.6924), while CatBoost and AdaBoost also showed competitive performance. This study demonstrates the potential of machine learning models in predicting DV risk using male demographic profiles. While acknowledging that findings derived from Nepal-specific data may not be directly generalizable to other sociocultural settings, the findings highlight critical socio-economic determinants such as education, wealth, and substance use and support the use of predictive modeling as a complementary tool for early identification and targeted intervention.
- Supplementary Content
- 10.26199/acu.8w440
- Jan 1, 2021
Twenty-five percent of Australian children are purported to have experiences of domestic and family violence. Despite this statistic, there is a lack of research in Australia with these children. To facilitate children’s engagement in domestic violence research, this study explored the barriers, enablers, and decision-making considerations of key gatekeepers and domestic violence researchers. In-depth, semi-structured interviews were held with 49 participants from five cohorts: domestic violence service providers, mothers with experiences of domestic/family violence, clinicians providing therapeutic interventions for children, Human Research Ethics Committee members, and domestic violence researchers. Themes about the barriers in domestic violence research with children concerned fears, safeguarding imperatives, and heightened risks. Domestic violence research with children was constructed as risky and dangerous. All cohorts, except domestic violence researchers, thought that this research could retraumatise children. The domestic violence service system and children being overshadowed in a closed adult-centric system emerged as further barriers in this research. Enablers in domestic violence research relate to the model and design of the research. Adopting a child-rights focus and trauma-safe methodology, along with having sector leadership, supportive gatekeepers and resources were identified as enablers. Attuned trauma-safe research, which is child-friendly, flexible, child-led, and creative, and which draws on the expertise of clinicians, further facilitates domestic violence research with children. To inform this research with children, an enabling model of attuned trauma-safe research, referred to as the STARR model, has been developed from the research findings.
- Book Chapter
6
- 10.4324/9781315612997-15
- Jan 2, 2018
This chapter draws on the existing domestic and sexual violence against older people research and make the links between three fields of inquiry: elder abuse; domestic violence; and sexual violence. Drawing on our own empirical research, it examines the evidence in relation to both victims and perpetrators of domestic and sexual violence and uses a number of case studies to highlight the overlaps and gaps in existing knowledge. The chapter focuses on older women, as the research shows they continue to be at increased risk of experiencing domestic or sexual violence compared to men. Furthermore, one of the major limitations with the existing domestic and sexual violence research is a lack of a unified definition of 'older'. An analysis of sexual violence against older women must therefore take into account the social structural position of older women in society and consider a range of factors, including age and gender, when exploring sexual violence against older people.
- Front Matter
23
- 10.1016/j.jpeds.2021.04.071
- May 5, 2021
- The Journal of pediatrics
Children Witnessing Domestic and Family Violence: A Widespread Occurrence during the Coronavirus Disease 2019 (COVID-19) Pandemic
- Research Article
9
- 10.13031/jnrae.15647
- Jan 1, 2023
- Journal of Natural Resources and Agricultural Ecosystems
Highlights Machine Learning (ML) models are identified, reviewed, and analyzed for HAB predictions. Data preprocessing is vital for efficient ML model development. ML models for toxin production and monitoring are limited. Abstract. Harmful algal blooms (HABs) are detrimental to livestock, humans, pets, the environment, and the global economy, which calls for a robust approach to their management. While process-based models can inform practitioners about HAB enabling conditions, they have inherent limitations in accurately predicting harmful algal blooms. To address these limitations, Machine Learning (ML) models can potentially leverage large volumes of IoT data to aid in near real-time predictions. ML models have evolved as efficient tools for understanding patterns and relationships between water quality parameters and HAB expansion. This review describes ML models currently used for predicting and forecasting HABs in freshwater ecosystems and presents model structures and their application for predicting algal parameters and related toxins. The review revealed that regression trees, random forest, Artificial Neural Network (ANN), Support Vector Regression (SVR), Long Short-Term Memory (LSTM), and Gated Recurrent Unit (GRU) are the most frequently used models for HABs monitoring. This review shows ML models' prowess in identifying significant variables influencing algal growth, HAB drivers, and multistep HAB prediction. Hybrid models also improve the prediction of algal-related parameters through improved optimization techniques and variable selection algorithms. While ML models often focus on algal biomass prediction, few studies apply ML models for toxin monitoring and prediction. This limitation can be associated with a lack of high-frequency toxin datasets for model development, and exploring this domain is encouraged. This review serves as a guide for policymakers and researchers to implement ML models for HAB prediction and reveals the potential of ML models for decision support and early prediction for HAB management. Keywords: Cyanobacteria, Freshwater, Harmful algal blooms, Machine learning, Water quality.
- Research Article
6
- 10.1186/s12886-025-04025-8
- Apr 10, 2025
- BMC Ophthalmology
BackgroundTo evaluate the effectiveness of machine learning (ML) models in predicting the occurrence of retinopathy of prematurity (ROP) and treatment need.MethodsFour ML models were created using 49 parameters known within the first 24 h post-birth and obtained during the initial screening examination, encompassing demographic, maternal, clinical, and neonatal intensive care unit-related data. The models’ performances were assessed using five machine learning (ML) classifier algorithms: logistic regression (LR), decision tree (DT), support vector machine (SVM), random forest (RF), and extreme gradient boosting (XGBoost). Performance metrics were calculated, and the top ten parameters with the highest predictive value were identified.ResultsIn the cohort of 355 preterm infants, Model I, predicting ROP development using birth data, achieved a balanced accuracy of 80%, with gestational age (GA), birth weight (BW) and mean corpuscular volume (MCV) as the top predictive parameters. Model II, predicting treatment-requiring ROP using birth data, exhibited a balanced accuracy of 81%. Key predictive parameters included low GA, BW, 1-minute and 5-minute APGAR scores, and low erythrocyte counts. For Model III, predicting ROP using the first screening examination data, and Model IV, predicting treatment-requiring ROP using the same data, the accuracy values were 80% and 66%, respectively, with BW, daily weight gain, total O2 support duration, and platelet/lymphocyte ratio emerged as the most significant predictive parameters in both models.ConclusionThis study demonstrates the potential of ML models to predict ROP development and treatment need. Incorporating clinical and intensive care-related parameters can enhance ROP screening and clinical decision-making.
- Research Article
9
- 10.1080/13632469.2025.2505974
- May 18, 2025
- Journal of Earthquake Engineering
The February 6, 2023, doublet earthquake in Kahramanmaraş, Türkiye, known for its multi-fault rupture, caused widespread destruction and affected numerous buildings across 17 provinces. This research focused on the utilization of machine learning (ML) models for the classification of an extensive earthquake-induced building damage dataset collected from detailed post-disaster inspections of 2,432,871 buildings, along with seismic activity data from nearby stations. The objective was to construct an ML-based model capable of accurately predicting the severity of earthquake-induced damage to buildings. The proposed model integrated the general characteristics of 965,270 reinforced concrete (RC) buildings to address the critical need for rapid and precise damage assessment following seismic events. Focusing on the robustness and reliability essential in complex seismic scenarios, this study evaluated nine ML models, including Decision Trees, Random Forest, XGBoost, and Logistic Regression. The evaluation of the performance metrics demonstrated the high predictive performance of Random Forest, achieving an accuracy of 93%. This model excelled in overall accuracy and performed comparatively better in the prediction of collapse instances. While KNN also achieved a similar accuracy and F1 score, it demonstrated lower performance across other metrics and faced limitations in accurately predicting building collapses. The results also demonstrated that using interpolated PGA data with the KDTree method can enhance the responsiveness of damage detection systems in the immediate aftermath of an earthquake. The findings provide insights into the optimization of model training within complex and large-scale datasets and establish the potential of ML models to enhance the efficiency and precision of post-seismic structural damage assessment.
- Research Article
4
- 10.2196/62805
- Feb 24, 2025
- Journal of medical Internet research
To address gaps in global understanding of cultural and social variations, this study used a high-performance machine learning (ML) model to predict adolescent substance use across three national datasets. This study aims to develop a generalizable predictive model for adolescent substance use using multinational datasets and ML. The study used the Korea Youth Risk Behavior Web-Based Survey (KYRBS) from South Korea (n=1,098,641) to train ML models. For external validation, we used the Youth Risk Behavior Survey (YRBS) from the United States (n=2,511,916) and Norwegian nationwide Ungdata surveys (Ungdata) from Norway (n=700,660). After developing various ML models, we evaluated the final model's performance using multiple metrics. We also assessed feature importance using traditional methods and further analyzed variable contributions through SHapley Additive exPlanation values. The study used nationwide adolescent datasets for ML model development and validation, analyzing data from 1,098,641 KYRBS adolescents, 2,511,916 YRBS participants, and 700,660 from Ungdata. The XGBoost model was the top performer on the KYRBS, achieving an area under receiver operating characteristiccurve (AUROC) score of 80.61% (95% CI 79.63-81.59) and precision of 30.42 (95% CI 28.65-32.16) with detailed analysis on sensitivity of 31.30 (95% CI 29.47-33.20), specificity of 99.16 (95% CI 99.12-99.20), accuracy of 98.36 (95% CI 98.31-98.42), balanced accuracy of 65.23 (95% CI 64.31-66.17), F1-score of 30.85 (95% CI 29.25-32.51), and area under precision-recall curve of 32.14 (95% CI 30.34-33.95). The model achieved an AUROC score of 79.30% and a precision of 68.37% on the YRBS dataset, while in external validation using the Ungdata dataset, it recorded an AUROC score of 76.39% and a precision of 12.74%. Feature importance and SHapley Additive exPlanation value analyses identified smoking status, BMI, suicidal ideation, alcohol consumption, and feelings of sadness and despair as key contributors to the risk of substance use, with smoking status emerging as the most influential factor. Based on multinational datasets from South Korea, the United States, and Norway, this study shows the potential of ML models, particularly the XGBoost model, in predicting adolescent substance use. These findings provide a solid basis for future research exploring additional influencing factors or developing targeted intervention strategies.
- Research Article
1
- 10.1136/jech-2024-222140
- Mar 12, 2025
- Journal of Epidemiology and Community Health
BackgroundViolence against women is a global problem with serious consequences. In response, Peru established women’s emergency centres or Centros de Emergencia Mujer (CEMs) in 1999, offering support services like psychological...
- Research Article
33
- 10.1016/j.agwat.2022.108115
- Dec 26, 2022
- Agricultural Water Management
A comparison of physical-based and machine learning modeling for soil salt dynamics in crop fields
- Research Article
4
- 10.3390/cancers16061199
- Mar 19, 2024
- Cancers
Simple SummaryEffective models for predicting non-functioning pituitary neuro-endocrine tumours (NF PitNET) recurrence and regrowth following surgical intervention remain elusive. Previous studies have identified conflicting risk factors in predictions of tumour progression for patients receiving surgery for NF PitNET. The aim of this study was to develop machine learning (ML) models to improve prediction of post-operative NF PitNET progression up to 15 years following surgery. ML models were shown to be effective for predicting tumour remission, stability, and regrowth, but were non-performant when predicting tumour recurrence or reduction in size. The extent of surgical resection was shown to have the strongest influence in the performant models, with lesser influence from age, tumour volume, and the use of post-operative radiotherapy, and with no influence shown from pre- or post-operative endocrine function.Post-operative tumour progression in patients with non-functioning pituitary neuroendocrine tumours is variable. The aim of this study was to use machine learning (ML) models to improve the prediction of post-operative outcomes in patients with NF PitNET. We studied data from 383 patients who underwent surgery with or without radiotherapy, with a follow-up period between 6 months and 15 years. ML models, including k-nearest neighbour (KNN), support vector machine (SVM), and decision tree, showed superior performance in predicting tumour progression when compared with parametric statistical modelling using logistic regression, with SVM achieving the highest performance. The strongest predictor of tumour progression was the extent of surgical resection, with patient age, tumour volume, and the use of radiotherapy also showing influence. No features showed an association with tumour recurrence following a complete resection. In conclusion, this study demonstrates the potential of ML models in predicting post-operative outcomes for patients with NF PitNET. Future work should look to include additional, more granular, multicentre data, including incorporating imaging and operative video data.
- Research Article
20
- 10.1016/j.spinee.2024.10.010
- Nov 4, 2024
- The Spine Journal
BACKGROUNDDysphonia is one of the more common complications following anterior cervical discectomy and fusion (ACDF). ACDF is the gold standard for treating degenerative cervical spine disorders, and identifying high-risk patients is therefore crucial. PURPOSEThis study aimed to evaluate different machine learning models to predict persistent dysphonia after ACDF. STUDY DESIGNA retrospective review of the nationwide Swedish spine registry (Swespine). PATIENT SAMPLEAll adults in the Swespine registry who underwent elective ACDF between 2006 and 2020. OUTCOME MEASURESThe primary outcome was self-reported dysphonia lasting at least 1 month after surgery. Predictive performance was assessed using discrimination and calibration metrics. METHODSPatients with missing dysphonia data at the 1-year follow-up were excluded. Data preprocessing involved one-hot encoding categorical variables, scaling continuous variables, and imputing missing values. Four machine learning models (logistic regression, random forest (RF), gradient boosting, K-nearest neighbor) were employed. The models were trained and tested using an 80:20 data split and 5-fold cross-validation, with performance metrics guiding the selection of the best model for predicting persistent dysphonia. RESULTSIn total, 2,708 were included in the study. Twelve key predictors were identified. Four machine learning models were tested, with the RF model achieving the best performance (AUC=0.794). The most significant predictors across models included preoperative NDI, EQ5Dindex, preoperative neurology, number of operated levels, and use of a fusion cage. The RF model, chosen for its superior performance, showed high sensitivity and consistent accuracy, but a low specificity and positive predictive value. CONCLUSIONSIn this study, machine learning models were employed to identify predictors of persistent dysphonia following ACDF. Among the models tested, the RF classifier demonstrated superior performance, with an AUC value of 0.790. The RF model identified NDI, EQ5Dindex, and number of fused vertebrae as key variables. These findings underscore the potential of machine learning models in identifying patients at increased risk for dysphonia persisting for more than 1 month after surgery.
- Research Article
7
- 10.2166/h2oj.2025.016
- Jul 1, 2025
- H2Open Journal
Reliable forecasting of extreme river water levels is crucial for flood mitigation, agricultural planning, and disaster preparedness, particularly in vulnerable regions like Bangladesh. Traditional hydrological models often struggle with the nonlinear dynamics of deltaic rivers influenced by monsoons, tides, and human activities. This study evaluates six regression-based machine learning (ML) models – linear regression (LR), random forest regression (RFR), XGBoost (XGBR), multilayer perceptron (MLPR), LightGBM (LGBMR), and polynomial regression (PR) – for predicting monthly maximum and minimum water levels in Bangladesh's Old Brahmaputra River. Using 34 years of data (1990–2024) from Islampur station, models were assessed via performance metrics and principal component analysis (PCA). Results demonstrated RFR's superior accuracy for maximum (R2 = 0.934, RMSE = 0.646 m) and minimum (R2 = 0.942, RMSE = 0.469 m) water levels, achieving the highest PCA-based composite scores. LGBMR and XGBR followed closely (R2 > 0.930, RMSE < 0.700 m), while MLPR performed poorly (RMSE = 0.978 m, R2 = 0.850). The significant performance gap between tree-based ensemble methods and other approaches highlights RFR's robustness in modeling nonlinear hydrological patterns. These findings underscore the potential of ML models for improving flood forecasting in data-scarce regions, aiding adaptive water management in the Brahmaputra Basin and similar flood-prone systems.
- Research Article
- 10.28933/irjph-2021-08-0106
- Jan 1, 2021
- International Research Journal of Public Health
Objectives: To determine the existence of a pattern of women most frequently victims of physical violence in Brazil over a period of 10 years. Methods: Data from the DATASUS platform were collected on the records of domestic, sexual and other violence, registered by physical violence against female persons between 2009 and 2018. Data from the Violence and Accident Surveillance System on characteristics of the violent act against women were also collected. The Brazilian Institute of Geography and Statistics was also used to collect data from the National Household Sample Survey (PNAD). For bibliographic reference, the descriptors “Domestic and Sexual Violence against Women”, “Domestic Violence” and “Domestic Violence” were searched on virtual data basis and Brazilian articles that were published within the period of the present study were included. Results: There is a continuous and rapid increase in the first half of the study period, with a slight deceleration between 2014 and 2016, followed by a new jump in records from 2017. As for race, the largest numbers are white women, 348428, and browns, 308902. Black women represent 68.25% of the total records of domestic, sexual and other violence, with 8.3% of the total records of physical violence. Conclusion: It is possible to estimate that black women are not making complaints or possibly are not being seen with due care to make them. As it is data that depends on denunciation, which is often not carried out, the results need consideration regarding assertiveness and reflection of reality.
- Research Article
6
- 10.1007/s10896-018-9985-0
- Aug 22, 2018
- Journal of Family Violence
The Engaging Men project aimed to identify facilitators, societal approaches to and support for domestic violence, and barriers to men’s participation in domestic violence research, assessing the importance of each factor. Participatory concept mapping was used with a convenience sample of men (n = 142) in person and online across Australia, Canada and the United States of America. Engaging Men identified 43 facilitators, societal approaches to and support for domestic violence, and/or barriers to men’s participation in domestic violence research. The strongest facilitators related to external connections, such as concern for women around them. Men also recognized societal approaches to and support for domestic violence and the strongest barriers centered on internal feelings, including fear, shame and guilt about being linked to domestic violence. This study suggests that providing a safe environment for men to express genuine thoughts, feeling and views about domestic violence is vital, yet rarely available in domestic violence research. Therefore, research opportunities need to be more effectively designed and incentivized to address challenging issues identified by men, such as fear, shame and guilt and offer meaningful opportunities to demonstrate positive change.
- Research Article
2
- 10.2151/jmsj.2025-018
- Jan 1, 2025
- Journal of the Meteorological Society of Japan. Ser. II
Tropical cyclones (TCs) are a threat to coastal regions in countries and areas situated in the tropics to, at times, mid-latitudes, and their threat is expected to escalate due to factors like global warming and urbanization. This emphasizes imperative need that warnings based on accurate and reliable forecasts be delivered to those who need them in order to prevent or mitigate TC impacts effectively. While conventional Numerical Weather Prediction (NWP) models have traditionally dominated TC forecasting at short to medium range lead times (i.e., up to two weeks), the emergence of Artificial Intelligence (AI) models, i.e., Machine Learning (ML) models trained on global reanalysis, has raised the possibility of such models competing and thus supplementing NWP models. Here, we examine the potential of ML models in operational TC forecasting, comparing them with conventional NWP models. The ML model used in this study is Pangu-Weather and TC forecasts by this ML model are compared with those from the operational global NWP model at the Japan Meteorological Agency, especially focusing on the track. All 64 named TCs for a period of 2021 to 2023 in the western North Pacific basin are verified. Results indicate that the ML forecasts exhibit smaller position errors compared to the NWP model, alleviate the westward bias around Japan, and retain its forecast accuracy for TCs with unusual paths, offering potential operational utility. Another benefit would be the ability to deliver forecast results to forecasters quicker than before, since the ML model’s forecast takes less than a minute. Meanwhile, challenges such as forecast bust cases and TC intensity, which are also present in NWP models, persist. A proposed way to utilize ML models at current operational systems would be to add ML-based track forecasts as one independent member of consensus forecasts.