A Simple Information Criterion for Variable Selection in High-Dimensional Regression.

Matthieu Pluntz,Cyril Dalmasso,Pascale Tubert-Bitter,Ismaïl Ahmed

doi:10.1002/sim.10275

Abstract

High-dimensional regression problems, for example with genomic or drug exposure data, typically involve automated selection of a sparse set of regressors. Penalized regression methods like the LASSO can deliver a family of candidate sparse models. To select one, there are criteria balancing log-likelihood and model size, the most common being AIC and BIC. These two methods do not take into account the implicit multiple testing performed when selecting variables in a high-dimensional regression, which makes them too liberal. We propose the extended AIC (EAIC), a new information criterion for sparse model selection in high-dimensional regressions. It allows for asymptotic FWER control when the candidate regressors are independent. It is based on a simple formula involving model log-likelihood, model size, the total number of candidate regressors, and the FWER target. In a simulation study over a wide range of linear and logistic regression settings, we assessed the variable selection performance of the EAIC and of other information criteria (including some that also use the number of candidate regressors: mBIC, mAIC, and EBIC) in conjunction with the LASSO. Our method controls the FWER in nearly all settings, in contrast to the AIC and BIC, which produce many false positives. We also illustrate it for the automated signal detection of adverse drug reactions on the French pharmacovigilance spontaneous reporting database.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

A Simple Information Criterion for Variable Selection in High-Dimensional Regression.

Abstract

Talk to us

Similar Papers

More From: Statistics in medicine

Lead the way for us

Journal: Statistics in medicine	Publication Date: Dec 12, 2024
License type: cc-by

Similar Papers

On GEE for Mean-Variance-Correlation Models: Variance Estimation and Model Selection.
Zhenyu Xu ... Jun Yan
Statistics in medicine | VOL. -
Zhenyu Xu, et. al.Zhenyu Xu ... Jun Yan
12 Dec 2024
Statistics in medicine | VOL. -

A Simple Information Criterion for Variable Selection in High-Dimensional Regression.
Matthieu Pluntz ... Ismaïl Ahmed
Statistics in medicine | VOL. -
Matthieu Pluntz, et. al.Matthieu Pluntz ... Ismaïl Ahmed
12 Dec 2024
Statistics in medicine | VOL. -

Linear Mixed Modeling of Federated Data When Only the Mean, Covariance, and Sample Size Are Available.
Marie Analiz April Limpoco ... Niel Hens
Statistics in medicine | VOL. -
Marie Analiz April Limpoco, et. al.Marie Analiz April Limpoco ... Niel Hens
11 Dec 2024
Statistics in medicine | VOL. -

-
-
Statistics in Medicine | VOL. -
--
10 Dec 2024
Statistics in Medicine | VOL. -

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Simple Information Criterion for Variable Selection in High-Dimensional Regression.

Abstract

Talk to us

Similar Papers

More From: Statistics in medicine