Computationally efficient whole-genome regression for quantitative and binary traits.
This study introduces REGENIE, a machine-learning method for efficient whole-genome regression of quantitative and binary traits, significantly reducing computational time and memory compared to existing methods, and demonstrating high accuracy and scalability on the UK Biobank dataset with over 400,000 individuals.
Genome-wide association analysis of cohorts with thousands of phenotypes is computationally expensive, particularly when accounting for sample relatedness or population structure. Here we present a novel machine-learning method called REGENIE for fitting a whole-genome regression model for quantitative and binary phenotypes that is substantially faster than alternatives in multi-trait analyses while maintaining statistical efficiency. The method naturally accommodates parallel analysis of multiple phenotypes and requires only local segments of the genotype matrix to be loaded in memory, in contrast to existing alternatives, which must load genome-wide matrices into memory. This results in substantial savings in compute time and memory usage. We introduce a fast, approximate Firth logistic regression test for unbalanced case-control phenotypes. The method is ideally suited to take advantage of distributed computing frameworks. We demonstrate the accuracy and computational benefits of this approach using the UK Biobank dataset with up to 407,746 individuals.
- Research Article
87
- 10.1016/j.ajhg.2008.12.005
- Jan 1, 2009
- American journal of human genetics
Fine Mapping on Chromosome 10q22-q23 Implicates Neuregulin 3 in Schizophrenia
- Research Article
1503
- 10.1038/s41588-018-0184-y
- Aug 13, 2018
- Nature genetics
In genome-wide association studies (GWAS) for thousands of phenotypes in large biobanks, most binary traits have substantially fewer cases than controls. Both of the widely used approaches, linear mixed model and the recently proposed logistic mixed model, perform poorly - producing large type I error rates - in the analysis of unbalanced case-control phenotypes. Here we propose a scalable and accurate generalized mixed model association test that uses the saddlepoint approximation to calibrate the distribution of score test statistics. This method, SAIGE, provides accurate p-values even when case-control ratios are extremely unbalanced. It utilizes state-of-art optimization strategies to reduce computational cost, and hence is applicable to GWAS for thousands of phenotypes by large biobanks. Through the analysis of UK Biobank data of 408,961 white British European-ancestry samples for >1400 binary phenotypes, we show that SAIGE can efficiently analyze large sample data, controlling for unbalanced case-control ratios and sample relatedness.
- Research Article
3
- 10.1186/1753-6561-5-s9-s73
- Nov 29, 2011
- BMC Proceedings
Clinical binary end-point traits are often governed by quantitative precursors. Hence it may be a prudent strategy to analyze a clinical end-point trait by considering a multivariate phenotype vector, possibly including both quantitative and qualitative phenotypes. A major statistical challenge lies in integrating the constituent phenotypes into a reduced univariate phenotype for association analyses. We assess the performances of certain reduced phenotypes using analysis of variance and a model-free quantile-based approach. We find that analysis of variance is more powerful than the quantile-based approach in detecting association, particularly for rare variants. We also find that using a principal component of the quantitative phenotypes and the residual of a logistic regression of the binary phenotype on the quantitative phenotypes may be an optimal method for integrating a binary phenotype with quantitative phenotypes to define a reduced univariate phenotype.
- Research Article
7
- 10.1002/gepi.22513
- Jan 24, 2023
- Genetic Epidemiology
In genome-wide association studies (GWAS) for thousands of phenotypes in biobanks, most binary phenotypes have substantially fewer cases than controls. Many widely used approaches for joint analysis of multiple phenotypes produce inflated type I error rates for such extremely unbalanced case-control phenotypes. In this research, we develop a method to jointly analyze multiple unbalanced case-control phenotypes to circumvent this issue. We first group multiple phenotypes into different clusters based on a hierarchical clustering method, then we merge phenotypes in each cluster into a single phenotype. In each cluster, we use the saddlepoint approximation to estimate the pvalue of an association test between the merged phenotype and a single nucleotide polymorphism(SNP) which eliminates the issue of inflated type I error rate of the test for extremely unbalanced case-control phenotypes. Finally, we use the Cauchy combination method to obtain an integrated pvalue for all clusters to test the association between multiple phenotypes and a SNP. We use extensive simulation studies to evaluate the performance of the proposed approach. The results show that the proposed approach can control type I error rate very well and is more powerful than other available methods. We also apply the proposed approach to phenotypes in category IX (diseases of the circulatory system) in the UK Biobank. We find that the proposed approach can identify more significant SNPs than the other viable methods we compared with.
- Research Article
471
- 10.1086/522036
- Dec 1, 2007
- American journal of human genetics
So Many Correlated Tests, So Little Time! Rapid Adjustment of P Values for Multiple Correlated Tests
- Dissertation
- 10.37099/mtu.dc.etdr/1429
- Jan 1, 2022
In genome-wide association studies (GWAS) for thousands of phenotypes in biobanks, most binary phenotypes have substantially fewer cases than controls. Many widely used approaches for joint analysis of multiple phenotypes in association studies produce inflated type I error rates for such extremely unbalanced case-control phenotypes. In our research, we develop two novel methods to jointly analyze multiple unbalanced case-control phenotypes to circumvent this issue. In the first method, we cluster multiple phenotypes into different clusters based on a hierarchical clustering method, then we merge phenotypes in each cluster into a single phenotype. In each cluster, we use the saddlepoint approximation to estimate the p-value of an association test between the merged phenotype and a SNP which eliminates the issue of inflated type I error rate of the test for extremely unbalanced case-control phenotypes. Finally, we use the Cauchy combination method to obtain an integrated p-value for all clusters to test the association between multiple phenotypes and a SNP. In the second method, we first construct a Multi-Layer Network (MLN) using all individuals with at least one case status among all phenotypes. Then, we introduce a computational efficient community detection method to group phenotypes into different disjoint clusters based on the MLN. The phenotypes in the same cluster are merged to a single phenotype which mainly eliminates the issue of inflated type I error rate of test for extremely unbalanced binary phenotypes. Finally, to test the association between all phenotypes and a SNP, we use the score test statistic to test the association between each merged phenotype and a SNP and then use the Omnibus test to obtain an overall p-value (MLN-O). Extensive simulation studies reveal that the newly proposed approaches can control type I error rates and are more powerful than other methods we compared with. The real data analyses also show that our methods outperform other methods we compared with.
- Research Article
43
- 10.1017/s0016672305007822
- Nov 25, 2005
- Genetical Research
Cluster analyses of gene expression data are usually conducted based on their associations with the phenotype of a particular disease. Many disease traits have a clearly defined binary phenotype (presence or absence), so that genes can be clustered based on the differences of expression levels between the two contrasting phenotypic groups. For example, cluster analysis based on binary phenotype has been successfully used in tumour research. Some complex diseases have phenotypes that vary in a continuous manner and the method developed for a binary trait is not immediately applicable to a continuous trait. However, understanding the role of gene expression in these complex traits is of fundamental importance. Therefore, it is necessary to develop a new statistical method to cluster expressed genes based on their association with a quantitative trait phenotype. We developed a model-based clustering method to classify genes based on their association with a continuous phenotype. We used a linear model to describe the relationship between gene expression and the phenotypic value. The model effects of the linear model (linear regression coefficients) represent the strength of the association. We assumed that the model effects of each gene follow a mixture of several multivariate Gaussian distributions. Parameter estimation and cluster assignment were accomplished via an Expectation-Maximization (EM) algorithm. The method was verified by analysing two simulated datasets, and further demonstrated using real data generated in a microarray experiment for the study of gene expression associated with Alzheimer's disease.
- Research Article
72
- 10.1007/s10519-011-9450-9
- Jan 1, 2011
- Behavior Genetics
The association of genetic variants with outcomes is usually assessed under an additive model, for example by the trend test. However, misspecification of the genetic model will lead to a reduction in power. More robust tests for association might therefore be preferred. A useful approach is to consider the maximum of the three test statistics under additive, dominant and recessive models (MAX3). The p-value however has to be adjusted to maintain the type I error rate. Previous studies and software on robust association tests have focused on binary traits without covariates. In this study we developed an analytic approach to robust association tests using MAX3, allowing for quantitative or binary traits as well as covariates. The p-values from our theoretical calculations match very well with those from a bootstrap resampling procedure. The methodology is implemented in the R package RobustSNP which is able to handle both small-scale studies and GWAS. The package and documentation are available at http://sites.google.com/site/honcheongso/software/robustsnp.
- Research Article
7
- 10.1038/s41588-025-02286-z
- Aug 11, 2025
- Nature genetics
Mixed-model association analysis (MMAA) is the preferred tool for performing genome-wide association studies. However, existing MMAA tools often have long runtimes and high memory requirements. Here we present LDAK-KVIK, an MMAA tool for analysis of quantitative and binary phenotypes. LDAK-KVIK is computationally efficient, requiring less than 10 CPU hours and 5 Gb memory to analyze genome-wide data for 350,000 individuals. Using simulated phenotypes, we show that LDAK-KVIK produces well-calibrated test statistics for both homogeneous and heterogeneous datasets. When applied to real phenotypes, LDAK-KVIK has the highest power among all tools considered. For example, across 40 quantitative UK Biobank phenotypes (average sample size 349,000), LDAK-KVIK finds 16% more independent, genome-wide significant loci than classical linear regression, whereas BOLT-LMM and REGENIE find 15% and 11% more, respectively. LDAK-KVIK can also be used to perform gene-based tests; across the 40 quantitative UK Biobank phenotypes, LDAK-KVIK finds 18% more significant genes than the leading existing tool. Last, LDAK-KVIK produces state-of-the-art polygenic scores.
- Research Article
16
- 10.1534/genetics.106.061275
- Nov 1, 2006
- Genetics
A novel method for Bayesian analysis of genetic heterogeneity and multilocus association in random population samples is presented. The method is valid for quantitative and binary traits as well as for multiallelic markers. In the method, individuals are stochastically assigned into two etiological groups that can have both their own, and possibly different, subsets of trait-associated (disease-predisposing) loci or alleles. The method is favorable especially in situations when etiological models are stratified by the factors that are unknown or went unmeasured, that is, if genetic heterogeneity is due to, for example, unknown genes x environment or genes x gene interactions. Additionally, a heterogeneity structure for the phenotype does not need to follow the structure of the general population; it can have a distinct selection history. The performance of the method is illustrated with simulated example of genes x environment interaction (quantitative trait with loosely linked markers) and compared to the results of single-group analysis in the presence of missing data. Additionally, example analyses with previously analyzed cystic fibrosis and type 2 diabetes data sets (binary traits with closely linked markers) are presented. The implementation (written in WinBUGS) is freely available for research purposes from http://www.rni.helsinki.fi/ approximately mjs/.
- Research Article
10
- 10.1038/sj.hdy.6800979
- Jun 6, 2007
- Heredity
Many binary phenotypes do not follow a classical Mendelian inheritance pattern. Interaction between genetic and environmental factors is thought to contribute to the incomplete penetrance phenomena often observed in these complex binary traits. Several two-locus models for penetrance have been proposed to aid the genetic dissection of binary traits. Such models assume linear genetic effects of both loci in different mathematical scales of penetrance, resembling the analytical framework of quantitative traits. However, changes in phenotypic scale are difficult to envisage in binary traits and limited genetic interpretation is extractable from current modeling of penetrance. To overcome this limitation, we derived an allelic penetrance approach that attributes incomplete penetrance to the stochastic expression of the alleles controlling the phenotype, the genetic background and environmental factors. We applied this approach to formulate dominance and recessiveness in a single diallelic locus and to model different genetic mechanisms for the joint action of two diallelic loci. We fit the models to data on the genetic susceptibility of mice following infections with Listeria monocytogenes and Plasmodium berghei. These models gain in genetic interpretation, because they specify the alleles that are responsible for the genetic (inter)action and their genetic nature (dominant or recessive), and predict genotypic combinations determining the phenotype. Further, we show via computer simulations that the proposed models produce penetrance patterns not captured by traditional two-locus models. This approach provides a new analysis framework for dissecting mechanisms of interlocus joint action in binary traits using genetic crosses.
- Research Article
1
- 10.3390/plants13172520
- Sep 7, 2024
- Plants (Basel, Switzerland)
Categorical (either binary or ordinal) quantitative traits are widely observed to measure count and resistance in plants. Unlike continuous traits, categorical traits often provide less detailed insights into genetic variation and possess a more complex underlying genetic architecture, which presents additional challenges for their genome-wide association studies. Meanwhile, methods designed for binary or continuous phenotypes are commonly used to inappropriately analyze ordinal traits, which leads to the loss of original phenotype information and the detection power of quantitative trait nucleotides (QTN). To address these issues, fast multi-locus ridge regression (FastRR), which was originally designed for continuous traits, is used to directly analyze binary or ordinal traits in this study. FastRR includes three stages of continuous transformation, variable reduction, and parameter estimation, and it can computationally handle categorical phenotype data instead of link functions introduced or methods inappropriately used. A series of simulation studies demonstrate that, compared with four other continuous or binary or ordinal approaches, including logistic regression, FarmCPU, FaST-LMM, and POLMM, the FastRR method outperforms in the detection of small-effect QTN, accuracy of estimated effect, and computation speed. We applied FastRR to 14 binary or ordinal phenotypes in the Arabidopsis real dataset and identified 479 significant loci and 76 known genes, at least seven times as many as detected by other algorithms. These findings underscore the potential of FastRR as a very useful tool for genome-wide association studies and novel gene mining of binary and ordinal traits.
- Research Article
12
- 10.1371/journal.pone.0207752
- Nov 21, 2018
- PLoS ONE
The logistic mixed model (LMM) is well-suited for the genome-wide association study (GWAS) of binary agronomic traits because it can include fixed and random effects that account for spurious associations. The recent implementation of a computationally efficient model fitting and testing approach now makes it practical to use the LMM to search for markers associated with such binary traits on a genome-wide scale. Therefore, the purpose of this work was to assess the applicability of the LMM for GWAS in crop diversity panels. We dichotomized three publicly available quantitative traits in a maize diversity panel and two quantitative traits in a sorghum diversity panel, and them performed a GWAS using both the LMM and the unified mixed linear model (MLM) on these dichotomized traits. Our results suggest that the LMM is capable of identifying statistically significant marker-trait associations in the same genomic regions highlighted in previous studies, and this ability is consistent across both diversity panels. We also show how subpopulation structure in the maize diversity panel can underscore the LMM’s superior control for spurious associations compared to the unified MLM. These results suggest that the LMM is a viable model to use for the GWAS of binary traits in crop diversity panels and we therefore encourage its broader implementation in the agronomic research community.
- Research Article
141
- 10.1016/j.ajhg.2018.12.012
- Jan 10, 2019
- The American Journal of Human Genetics
Efficient Variant Set Mixed Model Association Tests for Continuous and Binary Traits in Large-Scale Whole-Genome Sequencing Studies
- Preprint Article
4
- 10.1101/2024.07.25.24311005
- Jul 26, 2024
- medRxiv
ABSTRACTMixed-model association analysis (MMAA) is the preferred tool for performing a genome-wide association study, because it enables robust control of type 1 error and increased statistical power to detect trait-associated loci. However, existing MMAA tools often suffer from long runtimes and high memory requirements. We present LDAK-KVIK, a novel MMAA tool for analyzing quantitative and binary phenotypes. Using simulated phenotypes, we show that LDAK-KVIK produces well-calibrated test statistics, both for homogeneous and heterogeneous datasets. LDAK-KVIK is computationally-efficient, requiring less than ten CPU hours and 5Gb memory to analyse genome-wide data for 350k individuals. These demands are similar to those of REGENIE, one of the most efficient existing MMAA tools, and approximately ten times less than those of BOLT-LMM, currently the most powerful MMAA tool. When applied to real phenotypes, LDAK-KVIK has the highest power of all tools considered. For example, across the 40 quantitative UK Biobank phenotypes (average sample size 349k), LDAK-KVIK finds 16% more independent, genome-wide significant loci than classical linear regression, whereas BOLT-LMM and REGENIE find 15% and 11% more, respectively. LDAK-KVIK can also perform gene-based tests; across the 40 quantitative UK Biobank phenotypes, LDAK-KVIK finds 18% more significant genes than the leading existing tool. Lastly, LDAK-KVIK produces state-of-the-art polygenic scores.