Effect fusion using model-based clustering

Gertraud Malsiner-Walli,Helga Wagner,Daniela Pauger

doi:10.1177/1471082x17739058

Abstract

Abstract: In social and economic studies many of the collected variables are measured on a nominal scale, often with a large number of categories. The definition of categories can be ambiguous and different classification schemes using either a finer or a coarser grid are possible. Categorization has an impact when such a variable is included as covariate in a regression model: a too fine grid will result in imprecise estimates of the corresponding effects, whereas with a too coarse grid important effects will be missed, resulting in biased effect estimates and poor predictive performance. To achieve an automatic grouping of the levels of a categorical covariate with essentially the same effect, we adopt a Bayesian approach and specify the prior on the level effects as a location mixture of spiky Normal components. Model-based clustering of the effects during MCMC sampling allows to simultaneously detect categories which have essentially the same effect size and identify variables with no effect at all. Fusion of level effects is induced by a prior on the mixture weights which encourages empty components. The properties of this approach are investigated in simulation studies. Finally, the method is applied to analyse effects of high-dimensional categorical predictors on income in Austria.

Highlights

Researchers in medicine, social and economic sciences routinely collect data measured on a nominal scale as potential predictors in regression models
We proposed to specify a finite Normal mixture prior on the level effects of a categorical predictor to obtain a sparse representation of these effects
Level effects assigned to the same mixture component are fused, that is, their effects are replaced by the

Summary

Introduction

Researchers in medicine, social and economic sciences routinely collect data measured on a nominal scale as potential predictors in regression models. Spike and slab priors (Mitchell and Beauchamp, 1988; George and McCulloch, 1997; Ishwaran and Rao, 2005) in the Bayesian framework These methods are not appropriate for categorical covariates as only single level effects are selected or excluded from the model. By considering 22 markers, each with three levels, only a small number of levels is investigated Following this line of research we propose to achieve model-based clustering of level effects by specifying a finite Normal mixture prior. Specifying a sparsity inducing prior on the weights in an overfitting mixture avoids unnecessary splitting of superfluous components and encourages concentration of the posterior distribution on a sparse cluster solution and allows to estimate the number of effect groups from the data.

Effect clustering prior

Choice of hyperparameters

Posterior inference

MCMC sampling

Model-averaged estimates

Model selection

Simulation study

Set-up

Model selection results

Application

Discussion

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Statistical Modelling	Publication Date: Jan 15, 2018
Citations: 11	License type: cc-by

R Discovery Prime

R Discovery Prime

Effect fusion using model-based clustering

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Statistical Modelling

Lead the way for us

Similar Papers

A Renormalisation-Based Upscaling Technique for WAG Floods in Heterogeneous Reservoirs
P R King ... J W Barker
-
P R King, et. al.P R King ... J W Barker
12 Feb 1995
12 Feb 1995

Vorticity-based coarse grid generation for upscaling two-phase displacements in porous media
M.A Ashjari ... D Khoozan
Journal of petroleum science & engineering | VOL. 59
M.A Ashjari, et. al.M.A Ashjari ... D Khoozan
01 May 2007
Journal of petroleum science & engineering | VOL. 59

Calculation of Aerodynamic Noise for Centrifugal Fan of Air-Conditioner
Taku Iwase ... Hiroyasu Yoneyama
-
Taku Iwase, et. al.Taku Iwase ... Hiroyasu Yoneyama
07 Jul 2013
07 Jul 2013

Spatial+: A novel approach to spatial confounding.
Emiko Dupont ... Nicole H Augustin
Biometrics | VOL. 78
Emiko Dupont, et. al.Emiko Dupont ... Nicole H Augustin
30 Mar 2022
Biometrics | VOL. 78

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Effect fusion using model-based clustering

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Statistical Modelling