Predictive Monitoring of Wage-Band Classification in GOSI Data with Leakage Control and Out-of-Time Validation
Timely labor market monitoring is essential for policy design and operational planning, yet annual reports can mask turning points and subgroup heterogeneity. This paper develops a reproducible monitoring and prediction framework using administrative statistics from the General Organization for Social Insurance (GOSI) in the Saudi Open Data Portal. We document descriptive patterns in formal participation and insurable wages, including age-group dispersion, stable correlation structure, and explicit handling of an anomalous wage release and limited missing wage entries. We then formulate from non-salary administrative descriptors. Under leakage control, Random Forest models achieve accuracy around 0.71 across releases. Most errors are concentrated between adjacent wage bands, which is consistent with threshold discretization of a continuous wage distribution. To support operational deployment, we add out-of-time validation across releases and probabilistic assessment, showing that predictive skill transfers across updates and that calibration improves the reliability of probability scores for monitoring thresholds. Overall, the results indicate that administrative releases contain persistent actionable signals for wage segmentation without salary-derived inputs, supporting forecasting-oriented surveillance and early-warning dashboards.
- Book Chapter
- 10.1007/978-0-85729-410-4_361
- Jan 1, 2004
The present paper outlines how to assess the probability of PML (Probable Maximum Loss) caused by a fire or explosion in the oil industry. PML is assessed by use of both loss statistics and an ETA system developed by authors. In the study, at first, both types of statistics on big loss like PML and general loss are collected, and the probability of PML caused by a fire or explosion is estimated based on the global basis of the industry. Next, a model for probabilistic assessment of PML is developed using the ETA system, putting oil or gas leakage to an initial event. Successive events are detection of leakage, control of leakage and a breakout of fire or explosion. The probabilistic model makes it possible to estimate the fire or explosion risk for each fire or explosion scenario produced in an oil processing plant using failure rates regarding to the events given in the ETA system. Analyses of those fire and explosion scenarios using the proposed approach demonstrate that it presents consistent results with empirical facts. Thus, it is concluded that the system is feasible for the PML probabilistic assessment in the oil industry.
- Research Article
- 10.1016/j.mlwa.2026.100863
- Jun 1, 2026
- Machine Learning with Applications
Parallelized hybrid ensemble machine learning framework for scalable and accurate rainfall prediction
- Research Article
39
- 10.1016/j.forpol.2004.06.001
- Jul 24, 2004
- Forest Policy and Economics
Perceptions of communication in Norwegian forest management
- Research Article
70
- 10.1177/0143831x12448442
- Jun 7, 2012
- Economic and Industrial Democracy
Industrial relations scholarship has traditionally privileged union forms of employee participation. In more recent years there has been a shift to understand the role of participation in non-union firms. This article develops theory on employee participation through analysis of an Australian case study in the hotel sector. The authors find that formal participation mechanisms are useful and essential for both employees and managers, however formal participation leaves behind gaps which are partially filled with informal voice exchanges between employees and their managers.
- Research Article
59
- 10.1111/j.1469-8749.2009.03552.x
- Apr 13, 2010
- Developmental Medicine & Child Neurology
To determine patterns of participation and levels of enjoyment in young people with spinal cord injuries (SCI) and to assess how informal and formal participation varies across child, injury-related, household, and community variables. One hundred and ninety-four participants (106 males, 88 females; mean age 13y 2mo, SD 3y 8mo, range 6-18y) with SCI and their primary caregivers completed a demographics questionnaire and a standardized measure of participation (the Children's Assessment of Participation and Enjoyment, [CAPE]) at three pediatric SCI centers in a single hospital system in the United States. Their mean age at injury was 7 years 2 months (SD 5y 8mo, range 0-17y); 71% had paraplegia, and 58% had complete injuries. Young people participated more often in informal activities (t((174))=29.84, p<0.001) and reported higher enjoyment with these (t((174))=2.01, p=0.046). However, when engaging in formal activities, they participated with a more diverse group (t((174))=-16.26, p<0.001) and further from home (t((174))=-16.08, p<0.001). Aspects of informal participation were related to the child's age, sex, and injury level, and formal participation to the child's age and caregiver education. Caregiver education was more critical to formal participation among young people with tetraplegia than among those with paraplegia (F((4,151))=2.67, p=0.034). Points of intervention include providing more participation opportunities for young people with tetraplegia and giving caregivers the resources necessary to enhance their children's formal participation.
- Research Article
- 10.1149/10701.2967ecst
- Apr 24, 2022
- ECS Transactions
Data mining techniques have been used by several researchers to detect illnesses. Some methods are intended to predict a single sickness, while others are intended to predict a wide variety of diseases. It is also possible to improve the accuracy of sickness prediction. In this post, we provided an overview of data classification approaches that are available. These algorithms are mostly represented by themselves. The classification of data is a common and computationally difficult procedure. A massive amount of information must be analyzed in order to come up with a sound strategy for dealing with disease. There are a number of common scenarios, including as early disease detection, severity assessment, and prediction. This will decrease the progression of the disease, improve the well-being of the patients, and reduce the associated costs of medical care. In this regard, machine learning techniques have been employed. This article presents a machine learning based framework for heart disease data classification and prediction.
- Research Article
5
- 10.1016/j.egyr.2023.05.143
- Jun 6, 2023
- Energy Reports
Probabilistic load margin assessment considering forecast error of wind power generation
- Research Article
24
- 10.1016/j.ijfatigue.2024.108647
- Oct 16, 2024
- International Journal of Fatigue
Probabilistic framework for strain-based fatigue life prediction and uncertainty quantification using interpretable machine learning
- Research Article
4
- 10.1080/1351847x.2025.2476538
- Mar 15, 2025
- The European Journal of Finance
Predicting financial distress is vital for banks, regulators, and investors to manage risks and make informed decisions. However, the specific role of business plans in companies’ annual reports has received limited attention. This study introduces a new prediction framework that combines insights from business plans with financial and non-financial data. A key innovation of the framework is its weighting mechanism for tone variables, which links them directly to financial distress status. Using data from Chinese listed companies (2001–2022), the model shows improved predictive performance over short-term and long-term timeframes. Key indicators include words like ‘transformation,’ ‘low cost,’ and ‘enhancement,’ alongside measures of positive tone that capture optimistic language in business plans. The framework outperforms other advanced models, offering a practical tool for identifying financially troubled firms and supporting timely interventions.
- Research Article
13
- 10.1111/jgs.17282
- Jun 9, 2021
- Journal of the American Geriatrics Society
Older adults' susceptibility to mistreatment may be affected by their participation in social activities, but little is known about relationships between social participation and elder mistreatment. Cross-sectional analysis. National probability sample of older community-dwelling U.S. adults interviewed in 2015-2016, including 1268 women and 973 men (mean age 75 years and 76 years, respectively; 82% non-Hispanic white). Frequency of participation in formal activities (organized meetings, religious services, and volunteering) and informal social activities (visiting friends and family) was assessed by questionnaire. Elder mistreatment included emotional (four items), physical (two items), and financial mistreatment (two items) since age 60. Multivariable logistic regression examined associations between each type of social participation and elder mistreatment among men and women, adjusting for age, race/ethnicity, education, and comorbidity. Forty percent of women and 22% of men reported at least one form of mistreatment (emotional, physical, or financial). Women reporting at least monthly engagement in formal social activities were more likely to report emotional mistreatment (adjusted odds ratio (AOR) 1.59, 95% confidence interval (CI) 1.09-2.33). Among men, monthly organized meeting attendance was associated with increased odds of emotional mistreatment (AOR 1.34, 95% CI 1.01-1.93). Weekly informal socializing was inversely associated with emotional mistreatment (AOR 0.59, 95% CI 0.44-0.78) and financial mistreatment (AOR 0.59, 95% CI 0.42-0.85) among women. In this national cohort, older adults who were frequently engaged in formal social activities reported similar or higher levels of mistreatment than those with less frequent organized social participation. Older women with regular informal contact with family or friends were less likely to report some kinds of mistreatment. Strategies for detecting and mitigating elder mistreatment should consider differences in patterns of formal and informal social participation and their potential contribution to mistreatment risk.
- Research Article
1
- 10.5194/gmd-19-1027-2026
- Jan 30, 2026
- Geoscientific Model Development
Abstract. We propose a stochastic framework for wildfire spread prediction using deep generative diffusion models with ensemble sampling. In contrast to traditional deterministic approaches that struggle to capture the inherent uncertainty and variability of wildfire dynamics, our method generates probabilistic forecasts by sampling multiple plausible future scenarios conditioned on the same initial state. As a proof-of-concept, the model is trained on synthetic wildfire data generated by a probabilistic cellular automata simulator conditioned on canopy cover, vegetation density, and terrain slope for two real fires, namely the Chimney Fire in 2016 and the Ferguson Fire in 2018, both in California. To assess predictive performance and uncertainty representation under an identical neural network architecture, we compare a conventional supervised regression training paradigm against a conditional diffusion framework that employs ensemble sampling, and evaluate both approaches on independent ensemble test datasets. Across independent ensemble test sets, the diffusion surrogate consistently outperforms the deterministic baseline. It delivers lower errors in standard accuracy metrics such as mean squared error (MSE), exhibits higher spatial coherence as reflected by improved structural similarity index measure (SSIM) values, and generates samples of superior distributional quality according to the Fréchet inception distance (FID). Moreover, the diffusion-based model shows stronger probabilistic capability, as evidenced by higher scores in the hit rate (HR) metric, which we introduce as an uncertainty-aware verification measure. These results demonstrate that diffusion-based ensemble modelling provides a more flexible and effective approach for wildfire forecasting and, by capturing the distributional characteristics of future fire states, supports the generation of fire susceptibility maps that convey probabilistic risk information useful for assessment and operational planning in fire-prone environments.
- Research Article
- 10.3390/app16010157
- Dec 23, 2025
- Applied Sciences
This study presents an enhanced hybrid TLBO–ANN model for daily photovoltaic (PV) power generation prediction. By combining the strong nonlinear modeling capacity of Artificial Neural Networks (ANN) with the robust optimization capability of the Teaching–Learning-Based Optimization (TLBO) algorithm, the proposed framework effectively improves prediction accuracy and generalization performance. The model was trained using real meteorological and power generation data and validated on a grid-connected PV power plant in Türkiye. Results indicate that the hybrid TLBO–ANN approach outperforms the conventional ANN by achieving 39.97% and 37.46% improvements on the test subset and overall dataset, respectively. The improved convergence behavior and avoidance of local minima by TLBO contribute to this enhanced accuracy. Overall, the proposed hybrid model provides a powerful and practical tool for reliable PV power forecasting, which can facilitate better grid integration, operational planning, and energy management in renewable energy systems.
- Research Article
- 10.3390/jmse13122375
- Dec 15, 2025
- Journal of Marine Science and Engineering
Accurate wind information is essential for safe and efficient port operations, yet many small and medium-sized coastal ports lack dense meteorological instrumentation. This paper presents a data-driven framework for wind speed prediction at such ports by leveraging long-term historical measurements from nearby reference stations. Focusing on a real-world case study at the Chalkida port in Greece, the framework integrates both deterministic and Machine Learning (ML) models trained on historical wind patterns of archived wind data from four surrounding locations. We examine both short- and long-horizon prediction periods, using recently acquired wind measurements at the target port for model validation. Deterministic baselines include simple and weighted averaging schemes, while supervised ML methods, such as Multiple Linear Regression, Decision Trees, Support Vector Regression, Random Forests, and Gradient Boosting, are trained to capture complex spatiotemporal patterns. Experimental results highlight that ensemble-based ML models, particularly Gradient Boosting, achieve superior accuracy in short-term forecasting, while the optimal predictor varies with the forecast horizon. The proposed approach enables the deployment of virtual wind stations in data-scarce ports and can be periodically updated to dynamically select the most suitable model, thereby supporting climate adaptation strategies, localized wind monitoring, and operational planning without requiring dense local instrumentation.
- Research Article
3
- 10.1007/s10462-025-11140-x
- Mar 15, 2025
- Artificial Intelligence Review
Accurate forecasting of wind speed and direction is critical for the efficient integration of wind power into energy systems, ensuring reliable renewable energy production and grid stability. Traditional methods often struggle with capturing nonlinear interdependencies, quantifying uncertainties, and providing reliable long-term predictions, particularly in complex atmospheric conditions. To address these challenges, this study introduces multi-model Integration for dynamic forecasting (MIDF), an ensemble machine learning framework that combines the strengths of DeepAR and temporal fusion transformer (TFT) models through a two-step meta-learning process. MIDF leverages DeepAR’s probabilistic forecasting capabilities and TFT’s attention mechanisms to enhance accuracy, robustness, and interpretability. Using a custom meteorological dataset spanning January 2010 to May 2023, the model was evaluated against standalone alternatives across multiple metrics, including MSE, RMSE, and R2. MIDF achieved superior performance, with MSE, RMSE, and R2 values of 0.0035, 0.01913, and 0.89 for wind speed, and 0.00052, 0.02507, and 0.86 for wind direction, significantly reducing errors compared to existing methods. These results underscore the potential of ensemble learning in advancing wind forecasting accuracy, enabling more reliable renewable energy management, operational planning, and risk mitigation in meteorological applications.
- Research Article
20
- 10.1016/j.ijepes.2019.06.023
- Jun 14, 2019
- International Journal of Electrical Power & Energy Systems
Data-driven approach for spatiotemporal distribution prediction of fault events in power transmission systems