Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Robust Bayesian high-dimensional variable selection and inference with the horseshoe family of priors

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Robust Bayesian high-dimensional variable selection and inference with the horseshoe family of priors

Similar Papers
  • Research Article
  • Cite Count Icon 1
  • 10.14407/jrpr.2016.41.2.149
A Methodology for Estimating the Uncertainty in Model Parameters Applying the Robust Bayesian Inferences
  • Jun 30, 2016
  • Journal of Radiation Protection and Research
  • Joo Yeon Kim + 2 more

Background: Any real application of Bayesian inference must acknowledge that both prior distribution and likelihood function have only been specified as more or less convenient approximations to whatever the analyzer’s true belief might be. If the inferences from the Bayesian analysis are to be trusted, it is important to determine that they are robust to such variations of prior and likelihood as might also be consistent with the analyzer’s stated beliefs. Materials and Methods: The robust Bayesian inference was applied to atmospheric dispersion assessment using Gaussian plume model. The scopes of contaminations were specified as the uncertainties of distribution type and parametric variability. The probabilistic distribution of model parameters was assumed to be contaminated as the symmetric unimodal and unimodal distributions. The distribution of the sector-averaged relative concentrations was then calculated by applying the contaminated priors to the model parameters. Results and Discussion: The sector-averaged concentrations for stability class were compared by applying the symmetric unimodal and unimodal priors, respectively, as the contaminated one based on the class of e-contamination. Though e was assumed as 10%, the medians reflecting the symmetric unimodal priors were nearly approximated within 10% compared with ones reflecting the plausible ones. However, the medians reflecting the unimodal priors were approximated within 20% for a few downwind distances compared with ones reflecting the plausible ones. Conclusion: The robustness has been answered by estimating how the results of the Bayesian inferences are robust to reasonable variations of the plausible priors. From these robust inferences, it is reasonable to apply the symmetric unimodal priors for analyzing the robustness of the Bayesian inferences.

  • Research Article
  • Cite Count Icon 29
  • 10.1214/aos/1176348912
Invariance Properties of Density Ratio Priors
  • Dec 1, 1992
  • The Annals of Statistics
  • Larry Wasserman

Density ratio neighborhoods are classes of probabilities that are used in robust Bayesian inference. These classes are invariant under Bayesian updating and marginalization. This makes them computationally convenient in robust Bayesian inference. We show that this is the unique class of probabilities that has these invariance properties. Aside from its theoretical value, this result has computational implications as well.

  • Research Article
  • Cite Count Icon 45
  • 10.4172/2155-6180.s1-005
Bayesian Methods for High Dimensional Linear Models
  • Jan 1, 2013
  • Journal of Biometrics & Biostatistics
  • Himel Mallick Nengjun Yi

In this article, we present a selective overview of some recent developments in Bayesian model and variable selection methods for high dimensional linear models. While most of the reviews in literature are based on conventional methods, we focus on recently developed methods, which have proven to be successful in dealing with high dimensional variable selection. First, we give a brief overview of the traditional model selection methods (viz. Mallow's Cp, AIC, BIC, DIC), followed by a discussion on some recently developed methods (viz. EBIC, regularization), which have occupied the minds of many statisticians. Then, we review high dimensional Bayesian methods with a particular emphasis on Bayesian regularization methods, which have been used extensively in recent years. We conclude by briefly addressing the asymptotic behaviors of Bayesian variable selection methods for high dimensional linear models under different regularity conditions.

  • Research Article
  • Cite Count Icon 57
  • 10.1002/gepi.20353
Bayesian variable and model selection methods for genetic association studies
  • Jul 10, 2008
  • Genetic Epidemiology
  • Brooke L Fridley

Variable selection is growing in importance with the advent of high throughput genotyping methods requiring analysis of hundreds to thousands of single nucleotide polymorphisms (SNPs) and the increased interest in using these genetic studies to better understand common, complex diseases. Up to now, the standard approach has been to analyze the genotypes for each SNP individually to look for an association with a disease. Alternatively, combinations of SNPs or haplotypes are analyzed for association. Another added complication in studying complex diseases or phenotypes is that genetic risk for the disease is often due to multiple SNPs in various locations on the chromosome with small individual effects that may have a collectively large effect on the phenotype. Hence, multi-locus SNP models, as opposed to single SNP models, may better capture the true underlying genotypic-phenotypic relationship. Thus, innovative methods for determining which SNPs to include in the model are needed. The goal of this article is to describe several methods currently available for variable and model selection using Bayesian approaches and to illustrate their application for genetic association studies using both real and simulated candidate gene data for a complex disease. In particular, Bayesian model averaging (BMA), stochastic search variable selection (SSVS), and Bayesian variable selection (BVS) using a reversible jump Markov chain Monte Carlo (MCMC) for candidate gene association studies are illustrated using a study of age-related macular degeneration (AMD) and simulated data.

  • Research Article
  • 10.1002/sim.70173
Robust Bayesian Inference in the Multilevel Zero-Inflated Generalized Poisson Model.
  • Jul 1, 2025
  • Statistics in medicine
  • Mekuanint Simeneh Workie + 1 more

Outliers, over-dispersion, and zero inflation are issues with count data. Traditional models like Poisson and negative binomial often fail to account for these issues, leading to biased estimates and poor model fit. These frameworks are extended by the Zero-Inflated Generalized Poisson (ZIGP) model, which takes into consideration not only zero inflation but also over-dispersion or under-dispersion. However, in the presence of outliers and hierarchical data structures. This study develops a robust Bayesian inference framework for the multilevel ZIGP model. Standard Bayesian methods often lack robustness under model misspecification and in the presence of outlier data. The framework uses a Robust expectation solution (RES) algorithm and generalized Bayesian inference (GBI) for robust estimation against outliers. These approaches improve estimation accuracy using robust loss functions and scaling parameters to minimize the influence of outliers. Simulation studies confirm that the Robust Expectation Solution (RES) algorithm significantly outperformed the Expectation-Maximization (EM) algorithm in reducing bias and mean squared error (MSE), especially in the presence of outliers. Regular Bayesian and EM algorithms were more sensitive to outliers, leading to potential bias and instability in parameter estimates. Our robust Bayesian framework, specifically the Generalized Bayesian Inference (GBI), demonstrated improved robustness and stability under model misspecification and outlier contamination. The main results show that tuning quantiles and optimizing scaling parameters improved parameter calibration and reduced bias and mean square error (MSE). We applied the framework to neonatal mortality data, identifying key risk factors such as maternal education, wealth status, rural residence, and age at first birth.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 2
  • 10.3390/sym14020339
An Integrated Framework Based on GAN and RBI for Learning with Insufficient Datasets
  • Feb 7, 2022
  • Symmetry
  • Yao-San Lin + 2 more

Generative adversarial networks are known as being capable of outputting data that can imitate the input well. This characteristic has led the previous research to propose the WGAN_MTD model, which joins the common version of Generative Adversarial Networks and Mega-Trend-Diffusion methods. To prevent the data-driven model from becoming susceptible to small datasets with insufficient information, we introduced a robust Bayesian inference to the process of virtual sample generation based on the previous version and proposed its refined version, WGAN_MTD2. The new version allows users to append subjective information to the contaminated estimation of the unknown population, at a certain level. It helps Mega-Trend-Diffusion methods take into account not only the information from original small datasets but also the user’s subjective information when generating virtual samples. The flexible model will not be subject to the information from the present datasets. To verify the performance and confirm whether a robust Bayesian inference benefits the effective generation of virtual samples, we applied the proposed model to the learning task with three open data and conducted corresponding experiments for the significance tests. As the experimental study revealed, the integrated framework based on GAN and RBI, WGAN_MTD2, can perform better and lead to higher learning accuracies than the previous one. The results also confirm that a robust Bayesian inference can improve the information capturing from insufficient datasets.

  • Research Article
  • 10.1037/met0000837
How to apply Bayesian stochastic search variable selection with multiply imputed data.
  • Apr 16, 2026
  • Psychological methods
  • Sierra A Bainter + 2 more

Modern regularization and variable selection methods, such as least absolute shrinkage and selection operator (lasso) and Bayesian variable selection, are important tools for psychological researchers to reduce the risk of overfitting, improve prediction in future samples, and increase model interpretability. Although missing data are common in psychological data, it is not straightforward to combine principled methods for addressing missing data with these modern variable selection methods. This challenge is well illustrated in a recent article by Gunn et al. (2023) with a comparison of three approaches for combining lasso with multiple imputation to address missing data. Each of the surveyed approaches results in markedly different results in terms of predictors selected. Their findings underscore limitations of the lasso for the purpose of variable selection. In this article, we show how to implement a Bayesian variable selection method, stochastic search variable selection (SSVS), with multiply imputed data. SSVS is a principled and consistent method for variable selection, and we demonstrate advantages relative to lasso in an example data set and simulation study. It is straightforward to apply an ITS strategy for SSVS using existing software. (PsycInfo Database Record (c) 2026 APA, all rights reserved).

  • Research Article
  • Cite Count Icon 20
  • 10.1088/1475-7516/2023/10/021
Fast and robust Bayesian inference using Gaussian processes with GPry
  • Oct 1, 2023
  • Journal of Cosmology and Astroparticle Physics
  • Jonas El Gammal + 3 more

We present the GPry algorithm for fast Bayesian inference of general (non-Gaussian) posteriors with a moderate number of parameters. GPry does not need any pre-training, special hardware such as GPUs, and is intended as a drop-in replacement for traditional Monte Carlo methods for Bayesian inference. Our algorithm is based on generating a Gaussian Process surrogate model of the log-posterior, aided by a Support Vector Machine classifier that excludes extreme or non-finite values. An active learning scheme allows us to reduce the number of required posterior evaluations by two orders of magnitude compared to traditional Monte Carlo inference. Our algorithm allows for parallel evaluations of the posterior at optimal locations, further reducing wall-clock times. We significantly improve performance using properties of the posterior in our active learning scheme and for the definition of the GP prior. In particular we account for the expected dynamical range of the posterior in different dimensionalities. We test our model against a number of synthetic and cosmological examples. GPry outperforms traditional Monte Carlo methods when the evaluation time of the likelihood (or the calculation of theoretical observables) is of the order of seconds; for evaluation times of over a minute it can perform inference in days that would take months using traditional methods. GPry is distributed as an open source Python package (pip install gpry) and can also be found at https://github.com/jonaselgammal/GPry.

  • Research Article
  • Cite Count Icon 9
  • 10.1080/02664763.2018.1432576
Bayesian variable selection and coefficient estimation in heteroscedastic linear regression model
  • Feb 7, 2018
  • Journal of Applied Statistics
  • Taha Alshaybawee + 3 more

ABSTRACTIn many real applications, such as econometrics, biological sciences, radio-immunoassay, finance, and medicine, the usual assumption of constant error variance may be unrealistic. Ignoring heteroscedasticity (non-constant error variance), if it is present in the data, may lead to incorrect inferences and inefficient estimation. In this paper, a simple and effcient Gibbs sampling algorithm is proposed, based on a heteroscedastic linear regression model with an penalty. Then, a Bayesian stochastic search variable selection method is proposed for subset selection. Simulations and real data examples are used to compare the performance of the proposed methods with other existing methods. The results indicate that the proposal performs well in the simulations and real data examples. R code is available upon request.

  • Dissertation
  • 10.31274/td-20240329-105
Bayesian variable selection in ultra-high dimensional settings
  • Jan 1, 2021
  • Dongjin Li

This thesis is a collection of three papers focused on Bayesian variable selection for ultra-high dimensional problems. In various disciplines of modern scientific research, datasets are commonly available with tens of thousands of predictors but a limited number of observations. Nevertheless, only very few of these predictors are believed to be associated with the response, making variable selection for ultra-high dimensional setup an important and challenging problem. In this thesis, we present novel Bayesian methods for variable selection in ultra-high dimensional settings. In Chapter \ref{ch.sven}, we propose a Bayesian variable selection method, called SVEN, for Gaussian linear regression models. The method is based on a hierarchical Gaussian linear model with priors placed on the regression coefficients as well as on the model space. The use of degenerate {\it spike} priors on inactive variables and Gaussian {\it slab} priors for the important predictors results in sparsity of the regression coefficients and provides an analytically available form of the posterior probability of a model. The strong model selection consistency is shown to be attained when the number of predictors grows nearly exponentially with the sample size and even when the norm of mean effects solely due to the unimportant variables diverge, which is a novel attractive feature. We develop a scalable variable selection algorithm with an inbuilt screening method that efficiently explores the enormous model space, rapidly identify the regions of high posterior probabilities and make fast inference and prediction. To further mitigate multimodal posterior distributions, we use a temperature schedule whose values are guided by our model selection consistency results. An appealing byproduct of SVEN is the construction of novel model weight adjusted prediction intervals. To implement SVEN, we develop an R package called ``Bravo" which is now available on the Comprehensive R Archive Network (CRAN). In Chapter \ref{ch.bravo}, we describe the major features of the software and conduct step-by-step analysis on several examples to show the usages of the functions in the package. In Chapter \ref{ch.spsven}, we extend our SVEN model to include spatial random effects and call this Bayesian variable selection method SP-SVEN. Our SP-SVEN model is based on a hierarchical Gaussian linear mixed model where the well-known spike-and-slab priors are placed on the regression coefficients to achieve sparsity and a Gaussian intrinsic auto-regression prior is assigned to the spatial random effects. The use of Gaussian conjugate priors ensures the availability of the explicit form of the posterior distribution conditional on the two spatial parameters involved in the intrinsic auto-regression, and numerical integration is used to integrate out these two parameters to obtain the posterior distribution of a given model. For the priors on the two spatial parameters, we propose using the data from unsequenced varieties, if available, to build a hierarchical mixture model, and use the corresponding posterior distribution of the spatial parameters as their prior for the variable selection model. We also develop a scalable algorithm that embeds model based screening and uses fast Cholesky updates to compute the posterior probabilities, and thereby achieving fast exploration of the gigantic model space and rapid discovery of the high posterior regions. The outstanding performance of SVEN and SP-SVEN are demonstrated through a number of simulation studies and some real data examples from genome wide association studies in field trial experiments.

  • Research Article
  • 10.21914/anziamj.v55i0.7812
Bayesian variable selection and modelling for metastatic breast cancer data
  • Jun 21, 2014
  • ANZIAM Journal
  • Sarini Sarini + 2 more

A Bayesian model selection procedure is applied to data on 90 women with metastatic breast cancer. Protein covariates are measured on nucleus, cytoplasm, membrane, and stroma of primary breast carcinoma and lymph node metastasis tissue. Multiple imputation is performed to deal with missing data. Zellner's g-prior is used in the Bayesian variable selection procedure. The model space is reduced using posterior variable inclusion probabilities, and then posterior model probabilities are used to derive a candidate set of models. Bayesian model averaging is employed to robustly estimate survival time, and the goodness of fit of the derived model assessed by the correlation between estimated and observed survival times. The results show evidence of proteins having different rules in different parts of the tissue cell with respect to patient survival. Therefore, a recommendation is given on which part of the cell to observe certain proteins for prognosis. The models obtained are robust toward censoring and showed correlations between the observed and the predicted data between 0.7 and 0.84. References J. Adams, P. J. Carder, S. Downey, M. A. Forbes, K. MacLennan, V. Allgar, S. Kaufman, S. Hallam, R. Bicknell, J. J. Walker, F. Cairnduff, P. J. Selby, T. J. Perren, M. Lansdown, and R. E. Banks. Vascular endothelial growth factor (VEGF) in breast cancer: comparison of plasma, serum, and tissue VEGF and microvessel density and effects of tamoxifen. Cancer Res., 60(11):2898–2905, 2000. http://cancerres.aacrjournals.org/content/60/11/2898 . F. Bray, J.-S. Ren, E. Masuyer, and J. Ferlay. Global estimates of cancer prevalence for 27 sites in the adult population in 2008. Int. J. Cancer, 132(5):1133–1145, 2013. doi:10.1002/ijc.27711 . M. Clyde and E. I. George. Model uncertainty. Stat. Sci., 19(1):81–94, 2004. http://projecteuclid.org/download/pdfview_1/euclid.ss/1089808274 . A. R. T. Donders, G. J. M. G. van der Heijden, T. Stijnen, and K. G. M. Moons. Review: a gentle introduction to imputation of missing values. J. Clin. Epidemiol., 59(10):1087–1091, 2006. doi:10.1016/j.jclinepi.2006.01.014 . J. Ferlay, I. Soerjomataram, D. F. M. Ervik, F. Bray, R. Dikshit, S. Elser, C. Mathers, M. Rebelo, and D. M. Parkin. GLOBOCAN 2012: Estimated cancer incidence, mortality and prevalence worldwide in 2012. International Agency for Research on Cancer, World Health Organization, 2014. http://globocan.iarc.fr . E. I. George and R. E. McCulloch. Variable selection via Gibbs sampling. J. Am. Stat. Assoc., 88(423):881–889, 1993. doi:10.1080/01621459.1993.10476353 . E. I. George and R. E. McCulloch. Approaches for Bayesian variable selection. Stat. Sinica, 7(2):339–373, 1997. http://www3.stat.sinica.edu.tw/statistica/j7n2/j7n26/j7n26.htm . J. Geweke. Variable selection and model comparison in regression, volume 5 of Bayesian statistics, pages 609–620. Oxford University Press, 1996. J. A. Hoeting, D. Madigan, A. E. Raftery, and C. T. Volinsky. Bayesian model averaging: a tutorial. Stat. Sci., 14(4):382–401, 1999. http://www.jstor.org/stable/2676803 . D. G. Kleinbaum and M. Klein. Survival Analysis: A self-learning text. Statistics for Biology and Health. New York, Springer-Verlag, 2011. doi:10.1007/0-387-29150-4 . E. E. Leamer. Regression selection strategies and revealed priors. J. Am. Stat. Assoc., 73(363):580–587, 1978. doi:10.1080/01621459.1978.10480058 . E. E. Leamer. Specification searches: Ad hoc inference with nonexperimental data. Wiley New York, 1978. S. Mallett, P. Royston, R. Waters, S. Dutton, and D. G. Altman. Reporting performance of prognostic models in cancer: a review. BMC Med., 8(1):21, 2010. doi:10.1186/1741-7015-8-21 . H. C. McCosker. Prognostic significance of IGF and ECM induced signalling proteins in breast cancer patients. PhD thesis, School of Biomedical Sciences, QUT, 2012. http://eprints.qut.edu.au/53580/ . T. J. Mitchell and J. J. Beauchamp. Bayesian variable selection in linear regression. J. Am. Stat. Assoc., 83(404):1023–1032, 1988. doi:10.1080/01621459.1988.10478694 . K. Pantel and R. H. Brakenhoff. Dissecting the metastatic cascade. Nat. Rev. Cancer, 4(6):448–456, 2004. doi:10.1038/nrc1370 . R. Radpour, Z. Barekati, C. Kohler, W. Holzgreve, and X. Y. Zhong. New trends in molecular biomarker discovery for breast cancer. Genet. Test. Mol. Bioma., 13(5):565–571, 2009. doi:10.1089/gtmb.2009.0060 . A. E. Raftery, D. Madigan, and J. A. Hoeting. Bayesian model averaging for linear regression models. J. Am. Stat. Assoc., 92(437):179–191, 1997. doi:10.1080/01621459.1997.10473615 . A. Zellner. On assessing prior distributions and Bayesian regression analysis with g-prior distributions, volume 6 of Bayesian inference and decision techniques: Essays in Honor of Bruno De Finetti, pages 233–243. North-Holland, Amsterdam, The Netherlands, 1986. A. Zellner. An introduction to Bayesian inference in econometrics. New York: John Wiley and Sons, 1996.

  • Database
  • Cite Count Icon 3
  • 10.17863/cam.62808
BayesSUR: An R package package for high-dimensional multivariate Bayesian variable and covariance selection in linear regression
  • Apr 28, 2021
  • Apollo (University of Cambridge)
  • Leonardo Bottolo

In molecular biology, advances in high-throughput technologies have made it possible to study complex multivariate phenotypes and their simultaneous associations with high-dimensional genomic and other omics data, a problem that can be studied with high-dimensional multi-response regression, where the response variables are potentially highly correlated. To this purpose, we recently introduced several multivariate Bayesian variable and covariance selection models, e.g., Bayesian estimation methods for sparse seemingly unrelated regression for variable and covariance selection. Several variable selection priors have been implemented in this context, in particular the hotspot detection prior for latent variable inclusion indicators, which results in sparse variable selection for associations between predictors and multiple phenotypes. We also propose an alternative, which uses a Markov random field (MRF) prior for incorporating prior knowledge about the dependence structure of the inclusion indicators. Inference of Bayesian seemingly unrelated regression (SUR) by Markov chain Monte Carlo methods is made computationally feasible by factorisation of the covariance matrix amongst the response variables. In this paper we present BayesSUR, an R package, which allows the user to easily specify and run a range of different Bayesian SUR models, which have been implemented in C++ for computational efficiency. The R package allows the specification of the models in a modular way, where the user chooses the priors for variable selection and for covariance selection separately. We demonstrate the performance of sparse SUR models with the hotspot prior and spike-and-slab MRF prior on synthetic and real data sets representing eQTL or mQTL studies and in vitro anti-cancer drug screening studies as examples for typical applications.

  • Research Article
  • Cite Count Icon 10
  • 10.1109/tmi.2020.3031478
Robust Bayesian Analysis of Early-Stage Parkinson's Disease Progression Using DaTscan Images.
  • Feb 1, 2021
  • IEEE Transactions on Medical Imaging
  • Yuan Zhou + 2 more

This paper proposes a mixture of linear dynamical systems model for quantifying the heterogeneous progress of Parkinson's disease from DaTscan Images. The model is fitted to longitudinal DaTscans from the Parkinson's Progression Marker Initiative. Fitting is accomplished using robust Bayesian inference with collapsed Gibbs sampling. Bayesian inference reveals three image-based progression subtypes which differ in progression speeds as well as progression trajectories. The model reveals characteristic spatial progression patterns in the brain, each pattern associated with a time constant. These patterns can serve as disease progression markers. The subtypes also have different progression rates of clinical symptoms measured by MDS-UPDRS Part III scores.

  • Research Article
  • 10.7465/jkdi.2014.25.3.685
Robust Bayesian inference in finite population sampling with auxiliary information under balanced loss function
  • May 31, 2014
  • Journal of the Korean Data and Information Science Society
  • Eunyoung Kim + 1 more

In this paper, we develop Bayesian inference of the finite population mean with the assumption of posterior linearity rather than normality of the superpopulation in the presence of auxiliary information under the balanced loss function. We compare the performance of the optimal Bayes estimator under the balanced loss function with ones of the classical ratio estimator and the usual Bayes estimator in terms of the posterior expected losses, risks and Bayes risks.

  • Research Article
  • Cite Count Icon 12
  • 10.1086/psaprocbienmeetp.1994.1.193030
The Extent of Dilation of Sets of Probabilities and the Asymptotics of Robust Bayesian Inference
  • Jan 1, 1994
  • PSA: Proceedings of the Biennial Meeting of the Philosophy of Science Association
  • Timothy Herron + 2 more

We discuss two general issues concerning diverging sets of Bayesian (conditional) probabilities—divergence of “posteriors”—that can result with increasing evidence. Consider a set of probabilities typically, but not always, based on a set of Bayesian “priors.” Incorporating sets of probabilities, rather than relying on a single probability, is a useful way to provide a rigorous mathematical framework for studying sensitivity and robustness in Classical and Bayesian inference. See: Berger (1984, 1985, 1990); Lavine (1991); Huber and Strassen (1973); Walley (1991); and Wasserman and Kadane (1990). Also, sets of probabilities arise in group decision problems. See: Levi (1982); and Seidenfeld, Kadane, and Schervish (1989). Third, sets of probabilities are one consequence of weakening traditional axioms for uncertainty. See: Good (1952); Smith (1961); Kyburg (1961); Levi (1974); Fishburn (1986); Seidenfeld, Schervish, and Kadane (1990); and Walley (1991).

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant