A Bayesian Approach to Multiple-Output Quantile Regression Analysis under Informative Sampling
Abstract This article presents a Bayesian multiple-output quantile regression for complex survey data under informative sampling. Our approach relies on the asymmetric Laplace distributional assumption. From the location-scale mixture representation of this distribution, we introduce an Expectation–Maximization algorithm that provides a less computationally intensive alternative to the commonly used Markov Chain Monte Carlo algorithm for posterior inference. Our developments are mainly motivated by the joint analysis of growth indexes for Brazilian children under five.
- Research Article
23
- 10.22610/jsds.v4i4.751
- Apr 30, 2013
- Journal of Social and Development Sciences
This paper introduces Bayesian analysis and demonstrates its application to parameter estimation of the logistic regression via Markov Chain Monte Carlo (MCMC) algorithm. The Bayesian logistic regression estimation is compared with the classical logistic regression. Both the classical logistic regression and the Bayesian logistic regression suggest that higher per capita income is associated with free trade of countries. The results also show a reduction of standard errors associated with the coefficients obtained from the Bayesian analysis, thus bringing greater stability to the coefficients. It is concluded that Bayesian Markov Chain Monte Carlo algorithm offers an alternative framework for estimating the logistic regression model.
- Research Article
13
- 10.3847/1538-3881/ac8e66
- Oct 18, 2022
- The Astronomical Journal
Thanks to an enormous release of light curves of contact binaries, it is a challenge to derive the parameters of contact binaries using the Phoebe program and the Wilson–Devinney program with the Markov chain Monte Carlo (MCMC) algorithm. In this paper, we use neural network (NN) and MCMC algorithm to derive the parameters of contact binaries. The fitting of models is still done with the MCMC algorithm, but that the neural network is used to establish the mapping relationship between the parameters and the light curves generated beforehand by Phoebe. The NN model is trained with a set of Phoebe-generated light curves with known input parameters, and then combined with the MCMC algorithm to quickly obtain the posterior distribution of the parameters. Two NN models without and with the influence of third light are established, which can generate light curves with 100 points faster than Phoebe by about four orders of magnitude under the same running condition. In addition, the two models can generate the light curves with an error of less than a millimagnitude. The feasibility of NN and MCMC algorithm is also verified by the synthetic light curves generated by Phoebe and the light curves from Kepler survey data. NN and MCMC algorithms can quickly derive the parameters and the corresponding parameter errors of contact binaries from sky survey. These parameters can also be used as more precise initial input values for the objectives of individual detailed studies.
- Research Article
26
- 10.1007/s11336-014-9420-2
- Sep 1, 2015
- Psychometrika
The present paper proposes a hierarchical, multi-unidimensional two-parameter logistic item response theory (2PL-MUIRT) model extended for a large number of groups. The proposed model was motivated by a large-scale integrative data analysis (IDA) study which combined data (N = 24,336) from 24 independent alcohol intervention studies. IDA projects face unique challenges that are different from those encountered in individual studies, such as the need to establish a common scoring metric across studies and to handle missingness in the pooled data. To address these challenges, we developed a Markov chain Monte Carlo (MCMC) algorithm for a hierarchical 2PL-MUIRT model for multiple groups in which not only were the item parameters and latent traits estimated, but the means and covariance structures for multiple dimensions were also estimated across different groups. Compared to a few existing MCMC algorithms for multidimensional IRT models that constrain the item parameters to facilitate estimation of the covariance matrix, we adapted an MCMC algorithm so that we could directly estimate the correlation matrix for the anchor group without any constraints on the item parameters. The feasibility of the MCMC algorithm and the validity of the basic calibration procedure were examined using a simulation study. Results showed that model parameters could be adequately recovered, and estimated latent trait scores closely approximated true latent trait scores. The algorithm was then applied to analyze real data (69 items across 20 studies for 22,608 participants). The posterior predictive model check showed that the model fit all items well, and the correlations between the MCMC scores and original scores were overall quite high. An additional simulation study demonstrated robustness of the MCMC procedures in the context of the high proportion of missingness in data. The Bayesian hierarchical IRT model using the MCMC algorithms developed in the current study has the potential to be widely implemented for IDA studies or multi-site studies, and can be further refined to meet more complicated needs in applied research.
- Preprint Article
- 10.5194/egusphere-egu23-10326
- May 15, 2023
  Hydrologic models often are used to estimate streamflows at ungauged locations for infrastructure planning. These models can contain a multitude of parameters that themselves need to be estimated through calibration. Yet multiple sets of parameter values may perform nearly equally well in simulating flows at gauged sites, making these parameters highly uncertain. Markov Chain Monte Carlo (MCMC) algorithms can quantify parameter uncertainties; however, this can be computationally expensive for hydrological models. Thus, it is important to select an MCMC algorithm that is effective (converges to the true posterior parameter distribution), efficient (fast), reliable (consistent across random seeds) and controllable (insensitive to the algorithms hyperparameters). These characteristics can be assessed through algorithm diagnostics, but current MCMC diagnostics mostly focus on evaluating convergence of an individual search process, not diagnosing general problems of the algorithms. Therefore, additional diagnostics are required to represent algorithms sensitivity to their hyperparameters and to compare their performance across problems. Here, we propose new diagnostics to assess the effectiveness, efficiency, reliability and controllability of four MCMC algorithms: Adaptive Metropolis, Sequential Monte Carlo, Hamiltonian Monte Carlo, and DREAM(ZS). The diagnostic method builds off of diagnostics used to assess the performance of Multi-Objective Evolutionary Algorithms (MOEAs), and allows us to evaluate the sensitivity of the algorithms to their hyper-parameterization and compare their performance on multiple metrics, such as the Gelman-Rubin diagnostic and Wasserstein distance from the true posterior. We illustrate our diagnostics using the simple Hydrological Model (HYMOD) and several analytical test problems. This allows us to see which algorithms perform well on problems with different characteristics (e.g. known vs. unknown posterior shapes, uni- vs. multi-modality, low- vs. high-dimensionality). Since posterior shapes and modality are often unknown for hydrological problems, it is important to calibrate them with an MCMC algorithm that is robust across a wide variety of posterior shapes, and our new diagnostics allow for this identification.
- Research Article
41
- 10.1080/00949655.2011.603090
- Feb 1, 2013
- Journal of Statistical Computation and Simulation
Markov chain Monte Carlo (MCMC) algorithms have been shown to be useful for estimation of complex item response theory (IRT) models. Although an MCMC algorithm can be very useful, it also requires care in use and interpretation of results. In particular, MCMC algorithms generally make extensive use of priors on model parameters. In this paper, MCMC estimation is illustrated using a simple mixture IRT model, a mixture Rasch model (MRM), to demonstrate how the algorithm operates and how results may be affected by some commonly used priors. Priors on the probabilities of mixtures, label switching, model selection, metric anchoring, and implementation of the MCMC algorithm using WinBUGS are described, and their effects illustrated on parameter recovery in practical testing situations. In addition, an example is presented in which an MRM is fitted to a set of educational test data using the MCMC algorithm and a comparison is illustrated with results from three existing maximum likelihood estimation methods.
- Research Article
- 10.23977/acss.2016.11001
- Jan 1, 2016
- http://clausiuspress.com/journal/ACSS.html
In this paper we consider Bayesian analysis of the possible changes in hydrological time series by Markov chain Monte Carlo (MCMC) algorithm. We consider multiple change-points and various possible situations. The approach of Bayesian stochastic search selection is used for detecting and estimating the number and positions of possible change-point in a piecewise constant model. MCMC algorithm is used to estimate the posterior distributions of parameters. The result of the analysis is applied to the hydrological data sets of the major river net area of Shunde in China and the data set of Nile River. In order to further investigate the trends in each segment of the hydrological data sets, we consider the analysis of change-point regression model via MCMC algorithm.
- Supplementary Content
1
- 10.17635/lancaster/thesis/853
- Jul 18, 2020
- University of Lancaster
Epidemics often occur rapidly, with new cases being observed daily. Due to the frequently severe social and economic consequences of an outbreak, this is an area of research that benefits greatly from online inference. This motivates research into the construction of fast, adaptive methods for performing real-time statistical analysis of epidemic data. The aim of this thesis is to develop sequential Monte Carlo (SMC) methods for infectious disease outbreaks. These methods utilize the observed removal times of individuals, obtained throughout the outbreak. The SMC algorithm adaptively generates samples from the evolving posterior distribution, allowing for the real-time estimation of the parameters underpinning the outbreak. This is achieved by transforming the samples when new data arrives, so that they represent samples from the posterior distribution which incorporates all of the data. To assess the performance of the SMC algorithm we additionally develop a novel Markov chain Monte Carlo (MCMC) algorithm, utilising adaptive proposal schemes to improve its mixing. We test the SMC and MCMC algorithms on various simulated outbreaks, finding that the two methods produce comparable results in terms of parameter estimation and disease dynamics. However, due to the parallel nature of the SMC algorithm it is computationally much faster. The SMC and MCMC algorithms are applied to the 2001 UK Foot-and-Mouth outbreak: notable for its rapid spread and requirement of control measures to contain the outbreak. This presents an ideal candidate for real-time analysis. We find good agreement between the two methods, with the SMC algorithm again much quicker than the MCMC algorithm. Additionally, the performed inference matches well with previous work conducted on this data set. Overall, we find that the SMC algorithm developed is suitable for the real-time analysis of an epidemic and is highly competitive with the current gold-standard of MCMC methods, whilst being computationally much quicker.
- Research Article
30
- 10.1016/j.ecolmodel.2021.109608
- Jun 5, 2021
- Ecological Modelling
Sequential Monte-Carlo algorithms for Bayesian model calibration – A review and method comparison✰
- Research Article
30
- 10.1214/18-ba1100
- Mar 1, 2019
- Bayesian Analysis
Discovering temporal evolution of themes from a time-stamped collection of text poses a challenging statistical learning problem. Dynamic topic models offer a probabilistic modeling framework to decompose a corpus of text documents into “topics”, i.e., probability distributions over vocabulary terms, while simultaneously learning the temporal dynamics of the relative prevalence of these topics. We extend the dynamic topic model of Blei and Lafferty (2006) by fusing its multinomial factor model on topics with dynamic linear models that account for time trends and seasonality in topic prevalence. A Markov chain Monte Carlo (MCMC) algorithm that utilizes Pólya-Gamma data augmentation is developed for posterior sampling. Conditional independencies in the model and sampling are made explicit, and our MCMC algorithm is parallelized where possible to allow for inference in large corpora. Our model and inference algorithm are validated with multiple synthetic examples, and we consider the applied problem of modeling trends in real estate listings from the housing website Zillow. We demonstrate in synthetic examples that sharing information across documents is critical for accurately estimating document-specific topic proportions. Analysis of the Zillow corpus demonstrates that the method is able to learn seasonal patterns and locally linear trends in topic prevalence.
- Research Article
26
- 10.1080/15732479.2019.1628077
- Jun 25, 2019
- Structure and Infrastructure Engineering
Condition assessments of structures require prediction models such as empirical model and numerical simulation model. Generally, these prediction models have model parameters to be estimated from experimental data. Bayesian inference is the formal statistical framework to estimate the model parameters and their uncertainties. As a result, uncertainties associated with the model and measurement can be accounted for decision making. Markov Chain Monte Carlo (MCMC) algorithms have been widely employed. However, there still remain some implementation issues from the inappropriate selection of the proposal mechanism in Markov chain. Since the posterior density for a given problem is often problem-dependent and unknown, users require a trial-and-error approach to select and tune optimal proposal mechanism. To relieve this difficulty, various adaptive MCMC algorithms have been recently appeared. Users must understand their mechanism and limitations before applying the algorithms to their problems. However, there is no comprehensive work to provide detailed exposition and their performance comparison together. This study aims to bring together different adaptive MCMC algorithms with the goal of providing their mechanisms and evaluating their performances through comparative study. Three algorithms are chosen as the representative proposal mechanism. From comparative studies, the discussions were drawn in terms of performances, simplicity and computational costs for less-experienced users.
- Research Article
- 10.1080/10618600.2024.2410911
- Oct 7, 2024
- Journal of Computational and Graphical Statistics
Estimating the probability density of a population while preserving the privacy of individuals in that population is an important and challenging problem that has received considerable attention in recent years. While the previous literature focused on frequentist approaches, in this article, we propose a Bayesian nonparametric mixture model under differential privacy (DP) and present two Markov chain Monte Carlo (MCMC) algorithms for posterior inference. One is a marginal approach, resembling Neal’s algorithm 5 with a pseudo-marginal Metropolis-Hastings move, and the other is a conditional approach. Although our focus is primarily on local DP, we show that our MCMC algorithms can be easily extended to deal with global differential privacy mechanisms. Moreover, for some carefully chosen mechanisms and mixture kernels, we show how auxiliary parameters can be analytically marginalized, allowing standard MCMC algorithms (i.e., non-privatized, such as Neal’s Algorithm 2) to be efficiently employed. Our approach is general and applicable to any mixture model and privacy mechanism. In several simulations and a real case study, we discuss the performance of our algorithms and evaluate different privacy mechanisms proposed in the frequentist literature. Supplementary materials for this article are available online.
- Research Article
4
- 10.1007/s11222-020-09977-z
- Jan 1, 2021
- Statistics and Computing
We develop an Evolutionary Markov Chain Monte Carlo (EMCMC) algorithm for sampling spatial partitions that lie within a large and complex spatial state space. Our algorithm combines the advantages of evolutionary algorithms (EAs) as optimization heuristics for state space traversal and the theoretical convergence properties of Markov Chain Monte Carlo algorithms for sampling from unknown distributions. Local optimality information that is identified via a directed search by our optimization heuristic is used to adaptively update a Markov chain in a promising direction within the framework of a Multiple-Try Metropolis Markov Chain model that incorporates a generalized Metropolis-Hasting ratio. We further expand the reach of our EMCMC algorithm by harnessing the computational power afforded by massively parallel architecture through the integration of a parallel EA framework that guides Markov chains running in parallel.
- Research Article
51
- 10.1016/j.cageo.2018.01.011
- Feb 2, 2018
- Computers & Geosciences
This paper presents a new computer code developed to solve the 1D magnetotelluric (MT) inverse problem using a Bayesian trans-dimensional Markov chain Monte Carlo algorithm. MT data are sensitive to the depth-distribution of rock electric conductivity (or its reciprocal, resistivity). The solution provided is a probability distribution - the so-called posterior probability distribution (PPD) for the conductivity at depth, together with the PPD of the interface depths. The PPD is sampled via a reversible-jump Markov Chain Monte Carlo (rjMcMC) algorithm, using a modified Metropolis-Hastings (MH) rule to accept or discard candidate models along the chains. As the optimal parameterization for the inversion process is generally unknown a trans-dimensional approach is used to allow the dataset itself to indicate the most probable number of parameters needed to sample the PPD. The algorithm is tested against two simulated datasets and a set of MT data acquired in the Clare Basin (County Clare, Ireland). For the simulated datasets the correct number of conductive layers at depth and the associated electrical conductivity values is retrieved, together with reasonable estimates of the uncertainties on the investigated parameters. Results from the inversion of field measurements are compared with results obtained using a deterministic method and with well-log data from a nearby borehole. The PPD is in good agreement with the well-log data, showing as a main structure a high conductive layer associated with the Clare Shale formation.In this study, we demonstrate that our new code go beyond algorithms developend using a linear inversion scheme, as it can be used: (1) to by-pass the subjective choices in the 1D parameterizations, i.e. the number of horizontal layers in the 1D parameterization, and (2) to estimate realistic uncertainties on the retrieved parameters. The algorithm is implemented using a simple MPI approach, where independent chains run on isolated CPU, to take full advantage of parallel computer architectures. In case of a large number of data, a master/slave appoach can be used, where the master CPU samples the parameter space and the slave CPUs compute forward solutions.
- Research Article
20
- 10.1016/j.jcp.2016.11.024
- Dec 7, 2016
- Journal of Computational Physics
On an adaptive preconditioned Crank–Nicolson MCMC algorithm for infinite dimensional Bayesian inference
- Research Article
- 10.1002/sim.70198
- Jul 1, 2025
- Statistics in Medicine
ABSTRACTCorrelated survival data are prevalent in various clinical settings and have been extensively discussed in the literature. A common example is clustered survival data, where survival times are associated due to shared characteristics within clusters. In our study, we analyze invasive mechanical ventilation data collected from multiple intensive care units (ICUs) across Ontario, Canada. Patients within the same ICU exhibit similarities in clinical profiles and mechanical ventilation settings, leading to a correlation in their ventilation durations. To address this association, we introduce a shared frailty log‐logistic accelerated failure time model that accounts for intra‐cluster correlation through a cluster‐specific random intercept. We present a novel, fast variational Bayes (VB) algorithm for parameter inference and evaluate its performance using simulation studies varying the number of clusters and their sizes. We further compare the performance of our proposed VB algorithm with the h‐likelihood method and a Markov Chain Monte Carlo (MCMC) algorithm. The proposed algorithm delivers satisfactory results and demonstrates computational efficiency over the MCMC algorithm. We apply our method to ICU ventilation data from Ontario to investigate the ICU‐site random effect on ventilation duration.