Articles published on Finite mixture
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
3597 Search results
Sort by Recency
- Research Article
- 10.1016/j.cgh.2025.11.024
- Jul 1, 2026
- Clinical gastroenterology and hepatology : the official clinical practice journal of the American Gastroenterological Association
- Sylvester R Groen + 7 more
Delineating key determinants of health-related quality of life (HrQoL) in patients with disorders of the gut-brain-interaction (DGBI) remains challenging due to complex interplay of socioeconomic, psychological, and clinical factors. This exploratory comparative study aimed to identify factors and latent patient-profiles associated with generic and condition-specific HrQoL across different DGBI. Data from 4 clinical patient-cohorts were analyzed, including patients with functional dyspepsia ([FD]; n = 73), fecal incontinence ([FI]; n = 72), and irritable bowel syndrome ([IBS]; n = 419; 2 cohorts). Participants completed questionnaires on gastrointestinal (GI) symptoms, psychological factors, generic HrQoL (EQ-5D-5L), and condition-specific HrQoL. Finite mixture modeling identified latent clusters, and multivariate regression assessed factors associated with HrQoL outcomes. In patients with IBS and FD, higher depression scores and GI symptom severity were significantly associated with lower generic and condition-specific HrQoL (P < .001). In both FI and IBS cohorts, higher anxiety scores were associated with reduced HrQoL. Finite mixture modeling identified distinct latent clusters-2 in FD and FI cohorts, and 3 in IBS cohort-primarily defined by psychological comorbidity. Latent clusters with higher anxiety and depression scores showed markedly lower HrQoL outcomes. In contrast, socioeconomic and lifestyle factors appeared less relevant. There were no significant associations between GI symptom severity and HrQoL in FI. Psychological comorbidity appears the most salient factor associated with HrQoL, which was uniformly seen across different DGBI, whereas GI symptom severity was only associated with HrQoL in patients with FD and patients with IBS, but not in patients with FI. Although causality between psychological factors and HrQoL cannot be ascertained, our findings underscore the importance of holistic DGBI management that extends beyond symptom control to encompass the full spectrum of patient experience.
- Research Article
- 10.1016/j.jspi.2025.106369
- Jul 1, 2026
- Journal of Statistical Planning and Inference
- Guanfu Liu + 1 more
Homogeneity testing under finite mixtures of multivariate Poisson distributions
- Research Article
- 10.1080/03610926.2026.2686275
- Jun 9, 2026
- Communications in Statistics - Theory and Methods
- Andrea Nigri
We propose a novel spatial finite mixture model for the analysis and clustering of integer-valued differences between Poisson counts, motivated by applications where the scientific interest lies in net differences—such as mortality gaps or migration flows—rather than marginal event counts. Our approach models the observed differences using a mixture of Skellam distributions, providing a natural probabilistic framework for discrete, signed outcomes. Spatial dependence among regions is accommodated via a local Potts-type conditional specification on the latent cluster allocation, encouraging spatial coherence in the clustering structure. Parameter estimation is performed through maximum (pseudo) likelihood, via an EM-type algorithm through a mean-field approximation, with cluster assignment updates based on local maximization or mean-field approximations, and uncertainty quantified via bootstrap. Through simulation studies, we demonstrate the method’s ability to recover both true clustering and underlying parameters in scenarios with spatial structure and well-separated clusters. We further illustrate the approach using data on male-female neoplasms mortality gaps across Italian provinces, uncovering interpretable spatial patterns in excess mortality. The proposed methodology extends existing spatial mixture models to settings with integer-valued differences and provides a flexible tool for spatially structured analysis of comparative risks.
- Research Article
- 10.1080/15326349.2026.2682362
- Jun 5, 2026
- Stochastic Models
- Liang Jiao
Finite mixture models (FMMs) are widely applied in reliability engineering to characterize the lifetime distributions of heterogeneous populations. This paper establishes stochastic comparison results for two FMMs whose component distributions belong to the generalized Marshall–Olkin Topp Leone-G (MOTL-G) family, which uses a scale family as the baseline distribution. The analysis is conducted within the framework of majorization, chain majorization, and unordered majorization orders. We first derive sufficient conditions under which the usual stochastic order and the hazard rate order hold between two FMMs, in terms of majorization and chain majorization orders applied to the shape, scale, and mixing proportion parameters. Subsequently, we develop sufficient conditions based on unordered majorization orders for the usual stochastic order between two FMMs. Using aluminum alloy fatigue life data, we validate the flexibility of the proposed model and successfully identify its underlying heterogeneous structure. The theoretical findings are illustrated and validated through numerical examples and counterexamples.
- Research Article
1
- 10.1080/10618600.2026.2685040
- Jun 3, 2026
- Journal of Computational and Graphical Statistics
- Salvatore D Tomarchio + 2 more
Dimension-wise scaled normal mixtures (DSNMs) are a recently introduced class of continuous multivariate distributions that generalize the multivariate normal distribution by allowing (1) a broader form of central symmetry and (2) dimension-specific excess kurtosis. In this paper, we propose parsimonious finite mixtures of DSNMs for model-based clustering, specifically addressing scenarios with central symmetric clusters that differ in tail heaviness across dimensions. To achieve parsimony, we introduce structured constraints on the correlation, scale, and kurtosis parameters, resulting in a flexible family of 60 interpretable models. We outline expectation-maximization-based algorithms to obtain maximum likelihood estimates. As a concrete example, we focus on mixtures of dimension-wise scaled shifted exponential normal (DSSEN) distributions, a special case of DSNMs having closed-form joint density. Finally, we illustrate the practical advantages of the parsimonious DSSEN mixtures with applications to simulated and real-world data, benchmarking its performance against established mixtures of symmetric heavy-tailed distributions.
- Research Article
- 10.1007/s10985-026-09717-x
- Jun 2, 2026
- Lifetime data analysis
- Yingfa Xie + 3 more
Analyses of recurrent hypoglycemia are critical for effective treatment management in diabetic patients. Typically, within-subject dependency in such analyses is captured through subject-level frailty. Recent research has modeled recurrent hypoglycemia using the first hitting times of a reflected Brownian motion. A close examination of this approach reveals that it does not adequately account for varying frailties among individuals, which indicate notable heterogeneity. To address this gap, we propose a finite mixture model of the first hitting time distribution of the reflected Brownian motion. This model allows for component-specific regression coefficients and frailty parameters, providing nuanced insights into how risk factors differently affect patient subgroups. We employ a Bayesian framework for inference, utilizing Markov chain Monte Carlo for estimation. Model selection is conducted using the Deviance Information Criterion and the Logarithm of the Pseudo-Marginal Likelihood. The effectiveness of these criteria is assessed through simulation studies. Application to recurrent hypoglycemia modeling revealed two subgroups with different risk profiles, as reflected in their volatilities. Bayesian model comparison criteria favor the model with component-specific regression coefficients for volatilities. The subgroup with lower volatility exhibits a larger variance and, hence, a greater level of heterogeneity.
- Research Article
- 10.1515/ijb-2025-0066
- Jun 1, 2026
- The international journal of biostatistics
- Divan A Burger + 2 more
In right-skewed count data, the mean is disproportionately affected by a long upper tail, whereas the median remains a more representative measure of central tendency. Discrete Weibull (DW) regression links covariates to a shifted median, which in turn induces the exact integer median; however, a single DW component can fit poorly when the observed count distribution has a markedly heavier upper tail than a single-component model can accommodate. We propose a contaminated DW (cDW) regression that augments the baseline DW distribution with a more dispersed secondary component within a finite mixture while retaining a single shifted-median link. This mixture accommodates extreme counts more effectively, thereby stabilizing the median-based regression coefficients. The model accommodates general lower truncation at an arbitrary threshold c, including c = 1 for strictly positive outcomes and c = 0 for nonnegative counts, and is estimated using a straightforward Bayesian Markov chain Monte Carlo algorithm implemented in JAGS; R code accompanies the paper. Applied to hospital length-of-stay data, the cDW regression reduces the influence of outliers and achieves superior predictive performance relative to a single-component DW model, as demonstrated by leave-one-out cross-validation and a Kullback-Leibler influence diagnostic. Simulation experiments show that, under strongly heavy-tailed mixture settings, the cTDW model accurately recovers the regression coefficients and improves on the single-component TDW model. Because the added tail component can increase probability mass at both extremes, we further recommend embedding the cDW in a hurdle framework when structural zeros are present: the zero probability is modeled separately, and the heavy-tail mixture is applied only to positive counts. The cDW regression model provides a robust, median-centered alternative for analyzing skewed, possibly truncated count outcomes.
- Research Article
- 10.1016/j.bone.2026.117969
- Jun 1, 2026
- Bone
- Yi Zheng + 8 more
Proteomics-driven discovery of intervention windows and risk subtypes in osteoporosis: A prospective cohort study.
- Research Article
- 10.1016/j.sciaf.2026.e03284
- Jun 1, 2026
- Scientific African
- E Zinhom + 3 more
A nonparametric framework for linear–circular regression: Applications in environmental and biological sciences
- Research Article
- 10.1515/ijb-2025-0091
- May 28, 2026
- The international journal of biostatistics
- Wenxiao Zhou + 1 more
Missing data are common in real-world studies, yet the underlying missingness structure is often unknown, bringing additional uncertainty before an appropriate inference method can be applied. In this paper, we systematically examine two sources of such uncertainty: (1) the missing mechanism(s) involved and (2) the specific functional form of the missing model within a given mechanism. Focusing particularly on settings involving missing not at random (MNAR) data, we propose a two-step mixture-structure-based method, including a model filtering pre-screening step. The tasks of handling both sources of uncertainty and conducting reliable inference are then unified within a single EM-based framework. The core of our method lies in constructing a two-layer postulated mixture, which can be viewed as deliberately introducing an overfitted mixture - thereby enhancing flexibility and robustness to uncertainty. We consider two general scenarios in which the true data follow either a mixture or a non-mixture structure, and establish an identification framework for continuous finite mixtures potentially subject to MNAR. Simulation studies and a real-data application to the Medical Expenditure Panel Survey (MEPS) are utilized to demonstrate the performance of our method.
- Research Article
- 10.1080/08874417.2026.2668053
- May 7, 2026
- Journal of Computer Information Systems
- Ashish Varma + 1 more
ABSTRACT Artificial intelligence (AI) and other digital technologies, such as machine learning, data visualization, and workflow automation, have transformed audit tasks, procedures, and, to an extent, the entire audit profession by providing predictive and intelligent audit capabilities to support auditors. However, to benefit from AI, auditors require cognitive flexibility, which captures the extent to which the collaborating actors (here, audit team members) listen attentively without distraction and are open to multiple perspectives and possibilities, engage in the free flow of ideas and opinions based on their conversations with peers and clients, and evaluate different decision alternatives. This study aimed to ascertain how AI impacts audit team performance and what role cognitive flexibility plays in this process. Data were collected from 191 respondents from India, and partial least squares structural equation modeling (PLS-SEM) was used to conduct a comprehensive analysis of the conceptual model, including testing for mediation effects, testing for heterogeneity utilizing the finite mixture partial least squares (FIMIX) procedure and entropy values, testing for endogeneity using the Gaussian copula term, importance—performance map analysis (IPMA), and PLSpredict modeling. The results showed that cognitive flexibility is significant in AI-enabled audit engagements. This study concludes that AI generates transactional value and that auditors require proficiency in AI to perform audits effectively using such technologies.
- Research Article
- 10.1007/s43621-026-03081-4
- May 5, 2026
- Discover Sustainability
- Saralees Nadarajah + 2 more
Abstract This paper presents the first statistical modeling of peak times in maximum daily energy consumption for major industrial and service-sector companies in Tanzania, representing the first such analysis for any African country. Using high-resolution automated meter reading data from six of Tanzania’s largest energy consumers collected between 2020 and 2023, we model the time of peak demand as a circular variable. Due to the observed multimodality in peak times, a finite mixture of von Mises distributions is applied, with model parameters estimated via Markov chain Monte Carlo. The optimal number of mixture components varies considerably across companies-ranging from 4 to 10-reflecting diverse and company-specific demand patterns. Key circular statistics, including mean resultant, variance, skewness, and kurtosis, reveal substantial differences in peak concentration and temporal distribution: industrial firms exhibit more predictable, clustered peaks, while utility companies show highly variable and dispersed demand. Likelihood ratio tests confirm that peak-time distributions are stable across years, months, and days, indicating structurally embedded operational rhythms.
- Research Article
- 10.1093/biostatistics/kxag008
- May 4, 2026
- Biostatistics (Oxford, England)
- Emanuele Giorgi + 1 more
SummaryExisting approaches to modelling antibody concentration data are mostly based on finite mixture models that rely on the assumption that individuals can be divided into 2 distinct groups: seronegative and seropositive. Here, we challenge this dichotomous modelling assumption and propose a latent variable modelling framework in which the immune status of each individual is represented along a continuum of latent seroreactivity, ranging from minimal to strong immune activation. This formulation provides greater flexibility in capturing age-related changes in antibody distributions while preserving the full information content of quantitative measurements. We show that the proposed class of models can accommodate a large variety of model formulations, both mechanistic and regression-based, and also includes finite mixture models as a special case. We also propose a computationally efficient L_{2} -based estimator as an alternative to maximum likelihood estimation, which substantially reduces computational cost, and we establish its consistency. Through a case study on malaria serology, we demonstrate how the flexibility of the novel framework enables joint analyses across all ages while accounting for changes in transmission patterns. We conclude by outlining extensions of the proposed modelling framework and its relevance to other omics applications.
- Research Article
- 10.1080/03610918.2026.2666336
- May 2, 2026
- Communications in Statistics - Simulation and Computation
- Abdolnasser Sadeghkhani
Poisson mixture models provide a principled approach to overdispersed count data by representing latent heterogeneity through a random intensity. Beyond variance inflation, the mixing distribution governs extremal behavior, yet the tail implications of common observation mechanisms and mixture-selection uncertainty are often left implicit in applied analyses. Building on recent asymptotic tail classification results for Poisson mixtures and fast Bayesian model selection for finite mixtures, we develop a tail-aware inference framework for counts observed under imperfect detection and related recording constraints. We show how binomial thinning, zero truncation, and right censoring transform discrete tail functionals, yielding transparent links between mixing-tail classes and the observed extremes. Efficient posterior computation is achieved through collapsed sampling strategies that learn mixture complexity and regime structure. Simulations confirm the theoretical tail predictions across representative mixing families, and an analysis of replicated wildlife survey counts illustrates how the approach delivers interpretable tail diagnostics and predictive summaries for rare large counts.
- Research Article
- 10.22363/2658-4670-2026-34-1-24-39
- Apr 30, 2026
- Discrete and Continuous Models and Applied Computational Science
- Irina V Peshkova
In this paper the conditions to compare the extremal index of the stationary waiting time in the $M/G/1$ and $GI/M/1$ systems are obtained. These conditions include exponential asymptotic behaviour of waiting time tail and the order in failure rates for the interarrival intervals and for the service times in the systems to be compared. For $M/G/1$ system the obtained result is extended to the mixed service times with ordered components. If, in a $GI/G/1$ system, the service time is determined by a finite mixture whose dominant component of the equilibrium distribution belongs to the class of subexponential distributions then the tail of the limiting distribution of the stationary waiting time is equivalent to the tail of this distribution up to a constant obtained explicitly. Furthermore, the limiting distribution of the maximum of the stationary waiting time belongs to the maximum domain of attraction of the distribution of extreme values of the same type as the maximum of the random variables defined by the dominant component.
- Research Article
- 10.1108/josm-05-2025-0229
- Apr 29, 2026
- Journal of Service Management
- Mikèle Landry + 1 more
Purpose Service firms have the crucial responsibility of enabling value co-creation for their own and their customers' benefit. Therefore, it is imperative to gain a deeper understanding of how the organizational capabilities instilled in service interactions relate to the co-created value outcomes. Design/methodology/approach Grounded in configurational theory, organizational capabilities and value co-creation literature, this study aims to explain the complex relationships between interactional capabilities – a specific type of organizational capability that supports service interactions and contributes to a higher-level value co-creation capability – and key relational co-created value outcomes in two high-interaction service contexts. Relying on the service-dominant orientation model (SDO), the authors combine partial least squares structural equation modeling (PLS-SEM), finite mixture partial least squares (FIMIX-PLS) analysis and fuzzy set qualitative comparative analysis (fsQCA). Findings The PLS-SEM results establish the higher-level value co-creation capability construct and its positive relationship with the value outcomes. The FIMIX-PLS results indicate unobserved heterogeneity in the structural model, suggesting equifinality. The fsQCA results reveal how interactional capabilities, at a lower level, assemble in various configurations, resulting in both value co-creation and value no-creation. Originality/value Offering substantial theoretical and empirical reasoning with practical implications for service management, this study demonstrates how conjunction, equifinality and causal asymmetry characterize the relationships between interactional capabilities and value outcomes. In doing so, the paper expands knowledge about the value co-creation process in relation to value realization. The results also contribute to a more detailed understanding of the concept of value no-creation, as the absence of value realization.
- Research Article
- 10.1007/s41096-026-00290-y
- Apr 28, 2026
- Journal of the Indian Society for Probability and Statistics
- Smriti Sikha Sarma + 2 more
Wrapped Stable Distribution-based Inference for Finite Mixture Models: Principles and Applications
- Research Article
- 10.3390/su18094257
- Apr 24, 2026
- Sustainability
- Yu Tian + 1 more
Low-carbon energy transition (LET) has become an important global development strategy. However, in the contemporary industrial era, carbon emissions are intricately intertwined with economic growth based on the extensive use of fossil energy. To this end, the key to a more acceptable push for LET is to achieve carbon emissions decoupling (CED). The rapidly developing digital economy (DE) introduces novel possibilities for it. Using a Finite Mixture Model, this study aims to analyze how DE heterogeneously impacts CED across 66 countries from 2011 to 2022. As of 2022, 41% of countries attained strong decoupling status, 33% reached weak decoupling status. In terms of the effect of DE on CED, both chance and challenge are shown. DE exhibits dual effects: it enhances CED in high-education countries but hinders it in countries with rapid population growth. Government efficiency and gender equality amplify DE’s chance role, while natural gas or clean energy reliance weakens it. DE indirectly promotes CED via low-carbon behavior while raising risks through easier credit access. Meanwhile, the heterogeneity of institutional and economic characteristics in countries may influence the effect of DE on CED. These findings offer a theoretical foundation to reconcile economic sustainability with climate mitigation in digital transitions, providing actionable insights for policymakers to leverage DE’s potential in achieving SDG 13.
- Research Article
- 10.1177/1471082x261430008
- Apr 24, 2026
- Statistical Modelling
- Marco Alfò + 1 more
Longitudinal studies have known widespread use in the last years in several fields of research, as they allow to distinguish between different sources of variation. We may observe differences at the beginning of the study that stay persistent through time, and changes in the response that are due to temporal dynamics in the observed covariates. Individual-specific, time-constant, effects are often included in the linear predictor to allow for unobserved individual-specific, time constant, heterogeneity motivated by omitted individual features. The random effect approach to estimation is based on considering such effects as random variables, usually with a specific parametric distribution. This approach has been frequently criticized, as it is often employed not considering correlation between observed (i.e., covariates) and unobserved (i.e., random effects) terms. To solve this issue, we may explicitly account for correlation between observed and unobserved heterogeneity, using the so-called correlated effects approach. In this article, we show that a more general solution may be developed by estimating the random effect conditional distribution non-parametrically via a discrete probability distribution on a finite number of locations. The approach we propose is assessed via a large-scale simulation study and illustrated by the analysis of a benchmark dataset.
- Research Article
- 10.1038/s41467-026-71451-7
- Apr 15, 2026
- Nature communications
- Victor Yman + 23 more
Accurate serological tools are essential for monitoring the transmission of arboviruses with pandemic potential, yet cross-reactivity between closely related viruses hampers diagnostics and surveillance. Here, we develop a high-throughput multiplex serological assay to quantify antibody responses to 28 antigens from nine arboviruses (dengue, Zika, yellow fever, West Nile, Usutu, Japanese encephalitis, chikungunya (CHIKV), Mayaro (MAYV), and O'nyong-nyong virus) and apply it to over 4000 samples from epidemiologically distinct sites on four continents. We implement a flexible analytical method based on Bayesian finite mixture models and Receiver Operating Characteristic analysis to evaluate assay performance and define seropositivity thresholds. As a case study, we resolve cross-reactive and virus-specific responses for CHIKV and the emerging MAYV by combining competitive immunoassays with mathematical modelling of multiplex serological and epidemiological data. This approach yields cross-reactivity-adjusted estimates of local transmission dynamics, in agreement with existing epidemiological evidence, and reveals that CHIKV is more prone to induce cross-reactive antibody responses than MAYV. Our results demonstrate the power of combining multiplex serology with experimental validation and modelling to disentangle exposure histories in the face of serological cross-reactivity. This integrative approach holds promise for improving arbovirus surveillance, particularly in settings with overlapping transmission of multiple viruses and limited diagnostic capacity.