Articles published on Statistical learning
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
11434 Search results
Sort by Recency
- New
- Research Article
- 10.1016/j.cognition.2026.106512
- Jul 1, 2026
- Cognition
- Naama Schwartz + 5 more
Statistical learning performance is impacted by a previous learning experience: A predictive eye-movement study.
- New
- Research Article
- 10.1016/j.asoc.2026.115184
- Jul 1, 2026
- Applied Soft Computing
- Haodi Quan + 3 more
An explainable feature selection method through statistical reinforcement learning
- New
- Research Article
- 10.1109/tpami.2026.3672726
- Jul 1, 2026
- IEEE transactions on pattern analysis and machine intelligence
- Yan V G Ferreira + 4 more
Most machine learning methods assume fixed probability distributions, limiting their applicability in nonstationary real-world scenarios. While continual learning methods address this issue, current approaches often rely on closed-box models or require extensive user intervention for interpretability. We propose SyMPLER (Systems Modeling through Piecewise Linear Evolving Regression), an explainable model for time series forecasting in nonstationary environments based on dynamic piecewise-linear approximations. Unlike other locally linear models, SyMPLER uses generalization bounds from Statistical Learning Theory to automatically determine when to add new local models based on prediction errors, eliminating the need for explicit clustering of the data. Experiments show that SyMPLER can achieve comparable performance to both closed-box and existing explainable models while maintaining a human-interpretable structure that reveals insights about the system's behavior. In this sense, our approach conciliates accuracy and interpretability, offering a transparent and adaptive solution for forecasting nonstationary time series.
- New
- Research Article
- 10.47981/j.mijst.14(01)2026.568(99-107)
- Jun 30, 2026
- MIST INTERNATIONAL JOURNAL OF SCIENCE AND TECHNOLOGY
- Akinyemi Akinrotimi + 3 more
Pelvic Inflammatory Disease (PID) is an infection of the female reproductive organs, most frequently leading to problems such as infertility, chronic pelvic pain, and ectopic pregnancy. Due to its nonspecific symptoms, diagnosing PID remains problematic, often relying on clinical judgment and invasive procedures. This challenge highlights the need for more sensitive, non-invasive, and data-driven diagnostic tools to enhance early detection and treatment outcomes. As the symptoms are nonspecific, the diagnosis of PID remains problematic and needs the development of more sensitive and less invasive diagnostic tools. This study investigates combining Fisher's Exact Test for statistical feature selection with machine learning algorithms, including Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), and Artificial Neural Networks (ANN), to improve the diagnosis of Pelvic Inflammatory Disease (PID). The study analyses the clinical profiles of 200 patients diagnosed with PID and uses Fisher's Exact Test to identify key features associated with the condition's severity. These identified features were then used to train and evaluate the machine learning models. Among the algorithms used, the ANN performed best, achieving 91% accuracy, 90% precision, 92% recall, and 91% F1-score, demonstrating superior accuracy in predicting PID. The results show that the use of statistical and machine learning techniques, in their respective capacities, provides higher diagnostic accuracy, minimising misclassification errors often associated with traditional diagnostic methods. The study demonstrates the value of integrating data-driven approaches into clinical decision support systems to detect PID earlier and more accurately.
- New
- Research Article
- 10.1016/j.jenvman.2026.130304
- Jun 30, 2026
- Journal of environmental management
- Zecheng Li + 6 more
Machine learning early warning for urban heat risk with CMIP6 projections.
- New
- Research Article
- 10.29303/griya.v6i2.1250
- Jun 29, 2026
- Griya Journal of Mathematics Education and Application
- Nisa Husniatissibhi + 2 more
Numeracy skills are essential competencies that enable students to understand, process, and interpret mathematical information in various real-life contexts. However, students’ numeracy skills still need improvement, particularly in statistics learning. This study aimed to determine the effect of implementing the Problem-Based Learning (PBL) model through local culture-based Student Worksheets (LKPD) on the numeracy skills of tenth-grade students at SMAN 3 Mataram. This study employed a quasi-experimental method using a posttest-only control group design. The sample consisted of 78 students divided into an experimental class and a control class. Numeracy data were collected using a test instrument that had been validated by experts. The findings showed that students who learned through PBL with local culture-based LKPD achieved higher numeracy skills than those who received direct instruction. The results of the independent sample t-test indicated a significant difference between the two groups, with a very high level of instructional effect. These findings suggest that the implementation of PBL through local culture-based LKPD is effective in supporting the improvement of students’ numeracy skills in statistics learning.
- New
- Research Article
- 10.1186/s12874-026-02810-7
- Jun 29, 2026
- BMC medical research methodology
- Zahra Zolghadr + 6 more
Recent advances in data registration techniques have enabled researchers to analyze high-dimensional, low sample size (HDLSS) data, which poses challenges for traditional statistical and machine learning methods-especially in fields like neurodevelopmental disorder classification. Among various proposed solutions, the Data Maximum Dispersion Classifier (DMDC) has demonstrated superior performance. This study introduces two extensions: Sparse DMDC (SDMDC) and Longitudinal Sparse DMDC (LSDMDC), with the latter leveraging multi-way data frameworks. Simulated HDLSS datasets were generated from multivariate normal distributions with sparse structures. Evaluation of SDMDC involved varying class imbalance, predictor count, and correlation structures. Performance was measured via average F-measure and signal-to-noise ratio (SNR). Results showed that SDMDC outperformed other sparse models. Moreover, promising improvement was observed in model classification performance and the selection of effective predictors when the correlation between predictor variables was increased. In longitudinal simulations, LSDMDC maintained performance even as the number of predictors increased along the first dimension, effectively handling rank-1 and rank-2 data. Notably, LSDMDC achieved the best results under autoregressive correlation. In conclusion, SDMDC demonstrated clear advantages over other sparse classifiers, while LSDMDC proved highly effective for complex longitudinal structures.
- New
- Research Article
- 10.1080/07420528.2026.2693220
- Jun 29, 2026
- Chronobiology International
- Ecenur Özkul Erdoğan + 4 more
ABSTRACT This cross-sectional study aimed to investigate the combined effects of chronotype, Mediterranean diet adherence, and sleep quality on mental distress in university students, and to identify key predictors through advanced statistical modelling and machine learning approaches. A total of 659 undergraduate students aged 18–30 y in Konya, Türkiye, participated in the study. Chronotype was assessed using the Morningness – Eveningness Questionnaire (MEQ), sleep quality with the Pittsburgh Sleep Quality Index (PSQI), Mediterranean diet adherence with the Mediterranean Diet Adherence Screener (MEDAS), and mental distress with the Food-Mood Questionnaire (FMQ). Pearson correlation analyses, hierarchical regression, and path analysis were conducted to examine associations and potential mediation pathways. Additionally, an explainable XGBoost model with Shapley Additive Explanations (SHAP) values was applied to identify the most influential predictors of mental distress. Poorer sleep quality and evening chronotype were significantly associated with higher mental distress. Morning chronotype was positively associated with MEDAS but not with Body Mass Index (BMI). Hierarchical regression indicated that PSQI was the strongest predictor of mental distress, followed by chronotype. Path analysis further revealed that sleep quality significantly mediated the relationship between chronotype and mental distress, whereas Mediterranean diet adherence did not show a mediating effect. The XGBoost model demonstrated robust predictive performance, and SHAP analysis confirmed chronotype and sleep quality as the most influential variables in predicting mental distress. Evening chronotype and poorer sleep quality were associated with higher levels of mental distress in university students, while dietary quality did not show a significant independent association with mental distress in this sample. These findings highlight the importance of interventions focusing on circadian alignment and sleep hygiene to support mental well-being in young adults.
- New
- Research Article
- 10.1037/cep0000400
- Jun 25, 2026
- Canadian journal of experimental psychology = Revue canadienne de psychologie experimentale
- Shivang Shelat + 1 more
An abundance of evidence suggests that mind-wandering, or task-unrelated thought, poses significant detriments to performance. In 2013, however, Mooneyham and Schooler (2013) broadened mind-wandering's impact by reviewing emerging literature on both its costs and benefits. Research has now expanded enough that mind-wandering can no longer be treated as a simple failure of attention. In this sequel, we synthesize key findings from the past 13 years on when mind-wandering harms and when it helps. On the cost side, mind-wandering can reduce sensitivity to others' pain, spread lapses across peers in learning contexts, disrupt fine motor control, and weaken memory encoding. On the benefit side, mind-wandering is linked with improved memory consolidation, especially when thoughts return to recently learned material. Additional work suggests that mind-wandering can facilitate extraction of hidden probabilistic structure during statistical learning, support low-stakes decisions with less deliberative effort, and reduce semantic satiation during highly repetitive tasks. We also revisit mind-wandering's relation to mood, emphasizing that affective outcomes depend on more than just task-relatedness alone. This article outlines future directions to cultivate a sharper scientific understanding of how mind-wandering shapes different kinds of behavior. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
- New
- Research Article
- 10.1080/01431161.2026.2691978
- Jun 25, 2026
- International Journal of Remote Sensing
- Aniket Patel + 5 more
ABSTRACT Cloud base height (CBH) is a fundamental atmospheric parameter for weather forecasting, aviation safety, and climate research, as it provides key information on boundary layer structure, atmospheric stability, and cloud–radiation interactions. In this study, a two-step hybrid framework combining statistical cloud detection with machine learning (ML) regression for CBH estimation is developed and evaluated. In the first step, cloud presence is identified using a Variability Index (VI) defined as the ratio of the standard deviation of the backscatter profile to its peak value. This physically interpretable index shows strong class separability, with a large effect size (Cohen’s d ≈ 1.96), a maximum F1 score of about 0.83 at a VI of approximately 0.24, and an area under the ROC curve of 0.88, indicating effective cloud detection performance. In the second step, Multiple Linear Regression, Fine Tree Regression, Random Forest, and Gaussian Process Regression (GPR) models are applied to cloud-present profiles to estimate CBH. Among these, Random Forest, a tree based nonlinear ensemble model perform best, achieving a high correlation coefficients of about 0.94 ± 0.04. GPR, a kernel-based model, demonstrated slightly lower performance compared to Random Forest, achieving correlation coefficients of R = 0.91 ± 0.05, while the Fine Tree model showed the weakest performance among the nonlinear models tested in this study, achieving R = 0.89 ± 0.07. In contrast, Multiple Linear Regression model showed lowest accuracy with R = 0.58 ± 0.13. The results demonstrate that combining a simple, explainable statistical classification approach with advanced machine learning regression significantly improves the reliability and accuracy of CBH retrieval from Lidar backscatter data. The proposed framework is computationally efficient and shows strong potential for operational implementation in real-time atmospheric monitoring networks.
- New
- Research Article
- 10.1021/jasms.6c00119
- Jun 25, 2026
- Journal of the American Society for Mass Spectrometry
- Gustavo F Trindade + 15 more
Orbitrap secondary ion mass spectrometry (OrbiSIMS) combines high-resolution mass analysis with surface- and depth-resolved chemical imaging, enabling confident molecular identification across a wide range of materials and biological systems. As the number of OrbiSIMS instruments in operation continues to grow, establishing reproducible and comparable performance across laboratories has become increasingly important. Here, we report the results of a Versailles Project on Advanced Materials and Standards (VAMAS) interlaboratory comparison designed to assess Orbitrap noise characteristics, detection limits, and intensity scale calibration in OrbiSIMS. Using replicated spectra acquired from a common silver reference sample across ten instruments worldwide, we quantify the noise structure of the Orbitrap analyzer and determine scaling parameters that convert arbitrary intensity units into true ion counts. The results confirm that Orbitrap noise behavior is governed by the total ion population in the trap and exhibits consistent characteristics across instruments and Orbitrap models operated below saturation. The ratio between the calibration scaling parameter and the detector noise parameter (A/σ) is shown to be largely intensity-independent in the linear regime, providing a robust metric for comparing instrument performance. From the measured noise properties, detection limits are estimated for each instrument, highlighting the influence of ion load and thresholding on analytical sensitivity. In addition, instrument-specific weighted sum of Rician (WSoR) noise models are derived and shown to yield consistent scaling behavior, enabling noise-unbiased multivariate statistical analysis and machine learning. Together, these results establish a metrological framework for reproducible OrbiSIMS measurements, support meaningful cross-laboratory comparison, and provide a foundation for future studies addressing Orbitrap linearity and performance across the full dynamic range required for advanced SIMS imaging and depth-profiling applications. The concepts and metrics introduced are directly applicable to Orbitrap mass spectrometry more broadly and are a foundation for a robust framework that can be incorporated into QA/QC workflows.
- New
- Research Article
- 10.1038/s41598-026-58718-1
- Jun 24, 2026
- Scientific reports
- Šarlota Kaňuková + 6 more
Callus formation in Lavandula × intermedia varies widely depending on explant type, plant growth regulator composition, and cultivation duration, yet their combined effects remain insufficiently characterized. Here, 57 growth regulator treatments were evaluated using root- and stem-derived explants over a 15-week culture period. Callus induction occurred on most media and was typically initiated within the first three weeks, while biomass accumulation followed a biphasic pattern with a pronounced increase after week six, reaching up to 35g depending on treatment. A combination of 0.5mg L⁻¹ 2,4-D and 0.5mg L⁻¹ kinetin consistently produced the highest biomass. To model system behavior, five statistical and machine learning approaches were applied. XGBoost achieved the highest predictive accuracy on experimental data (R² ≈ 0.94), whereas Random Forest showed the most stable performance across independent validation datasets. Feature importance analysis identified culture duration as the dominant factor influencing biomass, while hormonal composition significantly affected both responses and explant type had only a minor contribution. Multi-objective optimization using NSGA-II revealed multiple high-performing solutions, while induction converged to a single near-optimal condition. These findings demonstrate that integrating experimental data with machine learning enables robust prediction and optimization of callus responses in Lavandula × intermedia.
- New
- Abstract
- 10.1093/oncolo/oyag205.002
- Jun 23, 2026
- The Oncologist
- Ghada Nouairia + 5 more
Background and AimsThere are currently no efficient tools for the diagnosis, prognosis, and monitoring of cholangiocarcinoma (CCA), even in high-risk populations such as patients with Primary Sclerosing Cholangitis (PSC). These tumors are rare, often diagnosed at a late stage with a poor survival rate. Current surveillance methods for early detection of CCA in PSC, using magnetic resonance imaging (MRI) and CA19-9 biomarker testing, have limited accuracy. Cancer-associated DNA methylation is changing during the cancer development. We hypothesize that such alterations are detectable in the peripheral blood of people a with CCA even at early disease stages and represent a promising non-invasive cancer detection method implementable clinically. In this study, we aimed to develop and validate a blood-derived DNA methylation-based early detection test of CCA.MethodsDNA was purified from the whole blood of 394 individuals including CCA, gallbladder cancer, PSC with and without CCA and healthy donors. Initially, a pilot cohort (n = 109) was analyzed using Illumina’s Infinum EPIC I array (850 K DNA methylation sites). Using statistical and machine learning methods, we identified differential DNA methylation sites associated with CCA as compared to PSC and healthy individuals.Subsequently, and thanks to funds from the Cholangiocarcinoma Foundation Fellowship program, we confirmed our findings in a larger cohort (n = 285), analyzed using Illumina’s Infinum EPIC II array (900 K sites). This validation cohort included external samples from Norway (n = 40) and longitudinal samples (n = 25) from people living with PSC with a follow-up of 5 to 13 years, with and without BTC development. Then, the validated methylation sites were used to build a machine learning model for accurate diagnosis of CCA. The efficacy of the test was assessed using area under the curve, sensitivity, specificity, root squared error, and precision recall.ResultsGenome-wide DNA methylation profiling of CCA revealed significant differentially methylated sites and regions related to CCA in PSC and non-PSC individuals. Around 2500 CCA-associated methylation sites were discovered on the pilot cohort then validated. First, a CCA early detection test in PSC and non-PSC individuals was developed using 100 DNA methylation sites. This non-invasive test offered greater accuracy (AUC = 0.97) than current surveillance methods and can be routinely repeated in high-risk individuals. Second, using enrichment analysis, we identified genes (e.g., GNAS), transcription factors, and pathways (e.g., MAPK, Ras and Rap1 signaling pathways) involved in CCA, that might have implications for disease monitoring.ConclusionOverall, this study shows that DNA methylation profiling of peripheral blood provides a reliable, non-invasive approach for both detecting CCA and monitoring disease progression, offering a promising tool for clinical implementation.
- New
- Research Article
- 10.1021/acs.jctc.6c00474
- Jun 23, 2026
- Journal of chemical theory and computation
- Yinkai Wu + 4 more
High-throughput computation is essential for the rational design of heterogeneous catalysts, and the quality of the initial adsorption configurations directly determines the efficiency and success rate of subsequent geometry optimization. Existing heuristic methods are constrained by rigid-body approximations, making it difficult for them to handle complex multidentate adsorption, and the low-precision force field correction schemes on which they rely lack generality. To address this issue, this work proposes a general algorithm, High-Throughput Initialization of Multidentate Adsorption Configurations (HiMac). The algorithm reformulates configuration generation as a multiobjective optimization problem. It employs forward kinematics to model molecular flexibility and combines it with a loss function based on parent molecule similarity, thereby unifying site selection with pose adjustment. A statistical learning module is further integrated to prioritize the exploration of chemically favorable sites and reduce the computational cost. Experiments show that HiMac generates high-quality initial configurations for geometry relaxation and transition-state searches, is applicable to arbitrary adsorbate-surface systems, and provides a general and efficient method to accelerate the rational design of heterogeneous catalysts.
- Research Article
- 10.1016/j.ijbiomac.2026.153148
- Jun 22, 2026
- International journal of biological macromolecules
- Henrique Solowej Medeiros Lopes + 9 more
Mechanical performance enhancement of cassava starch and açaí residue films through optimized hydroxypropylation reaction.
- Research Article
- 10.1038/s41598-026-59073-x
- Jun 21, 2026
- Scientific reports
- A Anwarsha + 1 more
Fault diagnosis of taper roller bearings needs to be accurate and efficient to ensure industrial machinery reliability. This research has developed a vibration-based fault diagnosis method which combines statistical feature ranking and machine learning classification. Vibration signals for the five different health conditions (healthy, inner race fault, outer race fault, roller fault, and cage fault) were collected from an SKF 32,206 taper roller bearing with changing speeds and loads. Two types of features, time-domain and frequency-domain, were derived and the features with the highest discriminating power were determined by one-way ANOVA and Kruskal Wallis statistical tests. Different combinations of features were tested with six classifiers: support vector machines (SVM), neural networks, discriminant analysis, naive bayes, decision trees, and nearest neighbor. It was found that Kruskal-Wallis feature-ranking benefits not only the result accuracy but also the computational efficiency, and the best feature set had 18 features. Thus, the linear SVM classifier yielded a classification accuracy of 99% and an AUC of 1, while requiring the least training time, demonstrating that it is suitable for real-time purposes. This work proposes a novel, dependable, and computationally efficient method of identifying faults in taper roller bearings, thus leading to greater automation of condition monitoring in industrial plants.
- Research Article
- 10.1016/j.concog.2026.104086
- Jun 19, 2026
- Consciousness and cognition
- Beat Meier + 1 more
Absolute pitch and sound-color synesthesia provide for unique learning opportunities.
- Research Article
- 10.1038/s41598-026-58459-1
- Jun 19, 2026
- Scientific reports
- Wesley J Wang + 1 more
Early-life adversity is widely linked to accelerated biological ageing, yet it remains unclear whether such associations reflect exposure during sensitive developmental periods, the cumulative burden of exposures, or temporal proximity to later outcomes. Here, we leverage life-history theory and a life course framework to nuance how the timing of adverse childhood experiences (ACEs) becomes biologically embedded through epigenetic ageing. Using longitudinal data from the Future of Families and Child Wellbeing Study (N=1,974), we apply statistical learning and structured life course modelling to test sensitive period, cumulative risk, and recency hypotheses across multiple domains of adversity (poverty, instability, deprivation, and maltreatment). We find that adversity exposure during specific developmental periods, rather than cumulative burden or recent exposure, are most strongly associated with epigenetic age acceleration in late childhood ([Formula: see text]=0.003). Moreover, the timing and direction of these effects vary by adversity type. Epigenetic ageing is in turn associated with later health-related risks ([Formula: see text]=0.29, SE=0.06; [Formula: see text]=1.62, SE=0.27) and demographic behaviour ([Formula: see text]=0.21, SE=0.08; [Formula: see text]=0.22, SE=0.11), and further mediates the association between ACEs and outcomes in young adulthood, particularly for BMI ([Formula: see text]=0.003, SE=0.002, [Formula: see text]=11%). These findings demonstrate that childhood adversity may be linked to biological ageing in developmentally specific and domain-dependent ways, with certain developmental periods appearing more sensitive to adversity exposure than others.
- Research Article
- 10.1016/j.dcn.2026.101767
- Jun 18, 2026
- Developmental cognitive neuroscience
- Tess Allegra Forest + 2 more
Neural correlates of learning speed reveal developmental differences in memory.
- Research Article
- 10.1038/s41598-026-56169-2
- Jun 17, 2026
- Scientific Reports
- Mahmoud Shabrawy + 3 more
The statistical prediction of seismic activity patterns from historical earthquake catalog data remains a major challenge in data-centered seismic hazard analysis because seismic time series are non-stationary, multi-scale, and clustered in nature. Existing data-driven seismic prediction pipelines often emphasize architectural innovation while giving less attention to systematic hyperparameter optimization, which is essential for achieving strong predictive performance. This work is motivated by the need for an integrated and computationally efficient data-driven time-series modeling framework. Accordingly, a hierarchical deep learning-metaheuristic optimization paradigm is proposed based on the Neural Hierarchical Interpolation for Time Series Forecasting (N-HITS) algorithm and the Gray Langurs Optimizer (GLO). We conduct a systematic benchmarking of N-HITS against state-of-the-art deep time-series prediction models trained under identical preprocessing and training conditions, followed by adaptive hyperparameter optimization. Baseline analysis showed that N-HITS, with a coefficient of determination (R^2) of 0.921 and a Mean Squared Error (MSE) of 0.00234, was the strongest standalone model. Following GLO-based hyperparameter optimization, performance improved to an R^2 of 0.9892 pm 0.0049 and an MSE of 7.980e-05 ± 7.980e-07, indicating substantial error reduction and higher convergence stability. These results highlight the importance of optimization intelligence in catalog-based statistical seismic activity prediction and position hierarchical deep learning with adaptive metaheuristic search as a scalable architecture for seismic trend monitoring. However, the proposed model relies only on historical seismic catalog patterns and does not incorporate tectonic processes or geophysical drivers; therefore, its outputs should be interpreted as statistical trend estimates rather than physically reliable earthquake predictions.