Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Improved polygenic prediction by Bayesian multiple regression on summary statistics

  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Accurate prediction of an individual’s phenotype from their DNA sequence is one of the great promises of genomics and precision medicine. We extend a powerful individual-level data Bayesian multiple regression model (BayesR) to one that utilises summary statistics from genome-wide association studies (GWAS), SBayesR. In simulation and cross-validation using 12 real traits and 1.1 million variants on 350,000 individuals from the UK Biobank, SBayesR improves prediction accuracy relative to commonly used state-of-the-art summary statistics methods at a fraction of the computational resources. Furthermore, using summary statistics for variants from the largest GWAS meta-analysis (n ≈ 700, 000) on height and BMI, we show that on average across traits and two independent data sets that SBayesR improves prediction R2 by 5.2% relative to LDpred and by 26.5% relative to clumping and p value thresholding.

Similar Papers
  • Research Article
  • Cite Count Icon 24
  • 10.1007/s12205-010-0087-7
Regional low flow frequency analysis using Bayesian regression and prediction at ungauged catchment in Korea
  • Sep 18, 2009
  • KSCE Journal of Civil Engineering
  • Sang Ug Kim + 1 more

Regional low flow frequency analysis using Bayesian regression and prediction at ungauged catchment in Korea

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 34
  • 10.1371/journal.pgen.1007856
Bayesian multiple logistic regression for case-control GWAS.
  • Dec 31, 2018
  • PLOS Genetics
  • Saikat Banerjee + 3 more

Genetic variants in genome-wide association studies (GWAS) are tested for disease association mostly using simple regression, one variant at a time. Standard approaches to improve power in detecting disease-associated SNPs use multiple regression with Bayesian variable selection in which a sparsity-enforcing prior on effect sizes is used to avoid overtraining and all effect sizes are integrated out for posterior inference. For binary traits, the logistic model has not yielded clear improvements over the linear model. For multi-SNP analysis, the logistic model required costly and technically challenging MCMC sampling to perform the integration. Here, we introduce the quasi-Laplace approximation to solve the integral and avoid MCMC sampling. We expect the logistic model to perform much better than multiple linear regression except when predicted disease risks are spread closely around 0.5, because only close to its inflection point can the logistic function be well approximated by a linear function. Indeed, in extensive benchmarks with simulated phenotypes and real genotypes, our Bayesian multiple LOgistic REgression method (B-LORE) showed considerable improvements (1) when regressing on many variants in multiple loci at heritabilities ≥ 0.4 and (2) for unbalanced case-control ratios. B-LORE also enables meta-analysis by approximating the likelihood functions of individual studies by multivariate normal distributions, using their means and covariance matrices as summary statistics. Our work should make sparse multiple logistic regression attractive also for other applications with binary target variables. B-LORE is freely available from: https://github.com/soedinglab/b-lore.

  • Peer Review Report
  • 10.7554/elife.69719.sa1
Decision letter: A proteome-wide genetic investigation identifies several SARS-CoV-2-exploited host targets of clinical relevance
  • Jun 28, 2021
  • John W Schoggins

Article Figures and data Abstract eLife digest Introduction Materials and methods Results Discussion Data availability References Decision letter Author response Article and author information Metrics Abstract Background: The virus SARS-CoV-2 can exploit biological vulnerabilities (e.g. host proteins) in susceptible hosts that predispose to the development of severe COVID-19. Methods: To identify host proteins that may contribute to the risk of severe COVID-19, we undertook proteome-wide genetic colocalisation tests, and polygenic (pan) and cis-Mendelian randomisation analyses leveraging publicly available protein and COVID-19 datasets. Results: Our analytic approach identified several known targets (e.g. ABO, OAS1), but also nominated new proteins such as soluble Fas (colocalisation probability >0.9, p=1 × 10-4), implicating Fas-mediated apoptosis as a potential target for COVID-19 risk. The polygenic (pan) and cis-Mendelian randomisation analyses showed consistent associations of genetically predicted ABO protein with several COVID-19 phenotypes. The ABO signal is highly pleiotropic, and a look-up of proteins associated with the ABO signal revealed that the strongest association was with soluble CD209. We demonstrated experimentally that CD209 directly interacts with the spike protein of SARS-CoV-2, suggesting a mechanism that could explain the ABO association with COVID-19. Conclusions: Our work provides a prioritised list of host targets potentially exploited by SARS-CoV-2 and is a precursor for further research on CD209 and FAS as therapeutically tractable targets for COVID-19. Funding: MAK, JSc, JH, AB, DO, MC, EMM, MG, ID were funded by Open Targets. J.Z. and T.R.G were funded by the UK Medical Research Council Integrative Epidemiology Unit (MC_UU_00011/4). JSh and GJW were funded by the Wellcome Trust Grant 206194. This research was funded in part by the Wellcome Trust [Grant 206194]. For the purpose of open access, the author has applied a CC BY public copyright licence to any Author Accepted Manuscript version arising from this submission. eLife digest Individuals who become infected with the virus that causes COVID-19 can experience a wide variety of symptoms. These can range from no symptoms or minor symptoms to severe illness and death. Key demographic factors, such as age, gender and race, are known to affect how susceptible an individual is to infection. However, molecular factors, such as unique gene mutations and gene expression levels can also have a major impact on patient responses by affecting the levels of proteins in the body. Proteins that are too abundant or too scarce may mean the difference between dying from or surviving COVID-19. Identifying the molecular factors in a host that affect how viruses can infect individuals, evade immune defences or trigger severe illness, could provide new ways to treat patients with COVID-19. Such factors are likely to remain constant, even when the virus mutates into new strains. Hence, insights would likely apply across all virus strains, including current strains, such as alpha and delta, and any new strains that may emerge in the future. Using such a 'natural experiment' approach, Karim et al. compared the genetic profiles of over 30,000 COVID-19 patients and a million healthy individuals. Nine proteins were found to have an impact on COVID-19 infection and disease severity. Four proteins were ranked as top priorities for potential treatment targets. One protein, called CD209 (also known as DC-SIGN), is involved in how the virus enters the host cells, and had one of the strongest associations with COVID-19. Two proteins, called IL-6R and FAS, were involved in the immune response and could be responsible for the immune over-activation often seen in severe COVID-19. Finally, one protein, called OAS1, formed part of the body's innate antiviral defence system and appeared to reduce susceptibility to COVID-19. Knowing more about the proteins that influence the severity of COVID-19 opens up new ways to predict, protect and treat patients who may have severe or fatal reactions to infection. Indeed, one of the identified proteins (IL-6R) had already been targeted in recent clinical trials with some encouraging results. Considering CD209 as a potential receptor for the virus could provide another avenue for therapeutics, similar to previously successful approaches to block the virus' known interaction with a receptor protein. Ultimately, this research could supply an entirely new set of treatment options to help combat the COVID-19 pandemic. Introduction At the current time, the coronavirus disease 2019 (COVID-19) pandemic is implicated in the deaths of more than 4 million people worldwide (Dong et al., 2020). Although effective vaccines have been developed to substantially reduce mortality and morbidity due to severe COVID-19, the emergence of mutated strains of the SARS-CoV-2 virus has challenged the effectiveness of existing vaccines and raised the urgency of identifying alternate therapeutic pathways to target the virus (Tegally, 2020; Erik et al., 2020 ; Collier et al., 2021). Nevertheless, it is likely that the mutated strains of SARS-CoV-2 will continue to exploit the same vulnerable host biology to bind onto and infect cells and, in susceptible individuals, evade immune defences and promote the excessive host inflammatory response that is characteristic of severe COVID-19 (Gordon et al., 2020a). Therefore, the identification of host proteins that play roles in COVID-19 susceptibility and severity remains crucial to the development of therapeutics as host protein mechanisms are independent of genomic mutations in the virus. An improved understanding of these therapeutically relevant virus-host pathways may also be important in combating viruses beyond SARS-CoV-2 (Perrin-Cocon et al., 2020). Several large-scale systematic experimental efforts have identified key host proteins that interact with viral proteins in the pathogenesis of severe COVID-19 (Gordon et al., 2020a; Gordon et al., 2020b; Bouhaddou et al., 2020). These notably include efforts to identify direct interactions with the spike protein of SARS-CoV-2, which mediates virus attachment onto receptors to infect host cells and is also the basis of most vaccines (Shang et al., 2020; Harvey et al., 2021). To complement in vitro host protein characterisation efforts, several groups have leveraged genetic datasets of human proteins and COVID-19 disease to identify therapeutically actionable candidate host proteins that are likely to play roles in enhancing COVID-19 susceptibility or to be involved in the pathogenesis of severe COVID-19 (Pairo-Castineira et al., 2021; Zhou et al., 2021). One of the approaches used was Mendelian randomisation (MR). MR simulates the design of randomised trials, with the underlying principle that randomisation of alleles at conception offers the opportunity to examine approximate differences in average risk of disease between comparable groups in a population that differ only in the distribution of the risk factor of interest (Davies et al., 2018), for example, protein abundance (Zheng et al., 2020). This allows the use of alleles as genetic instruments representing genetically predicted protein levels to proxy effects of pharmacological modulation of the protein. Some of the clinically actionable proteins identified by the MR approach are part of type I interferon signalling (encoded by genes: IFNAR2, TYK2, OAS1) and interleukin-6 (IL-6) signalling pathways (IL6R). Only one of these proteins (encoded by OAS1) had any evidence of genetic colocalisation, that is, evidence that genetic associations of the protein and COVID outcomes shared the same causal genetic signal (Zhou et al., 2021). An additional protein that was supported by both MR and genetic colocalisation tests was ABO (Zhou et al., 2021), reported in several published genome-wide association studies (GWAS) of COVID-19 (Pairo-Castineira et al., 2021; Ellinghaus et al., 2020). In response to the first published GWAS of COVID-19, we reported findings that link the ABO signal with a number of clinically actionable targets including coagulation factors (von Willebrand factor [vWF], and Factor VIII [F8]), IL-6, and CD209/DC-SIGN (Karim et al., 2020). However, in most of the previous MR studies (Pairo-Castineira et al., 2021; Zhou et al., 2021), investigators only used curated cis-acting variants (genetic variants near or in the gene encoding the relevant protein) as genetic instruments to represent effects of genetically predicted protein concentrations, rather than genome-wide instruments. While the use of cis-acting variants can minimise the risk of horizontal pleiotropic effects (i.e. associations driven by other proteins not on the causal pathway for the disease), it can suffer from lower power than a genome-wide analysis due to fewer available instruments (Zheng et al., 2020). Furthermore, in previous protein-COVID-19 MR studies, genetic colocalisation tests were carried out only for protein-phenotype associations that were significant in the MR analysis, potentially excluding many protein-phenotype associations that may share the same causal genetic signal but are underpowered in a proteome-wide MR approach. In the present study, we expanded on these previous reports by undertaking a proteome-wide two-sample pan- and cis-MR analysis using the Sun et al. GWAS (Sun et al., 2018) of plasma protein concentrations and several COVID-19 GWAS phenotypes from the ICDA COVID-19 Host Genetics Initiative (October 2020 release) (Huang et al., 2020). First, we showed that genetically predicted circulating ABO protein was associated with COVID-19 susceptibility and severity and the lead ABO signal was associated strongly with plasma concentrations of soluble CD209. Second, we collected evidence for a direct mechanism of interaction between the SARS-CoV-2 spike protein and human CD209 protein. Third, we performed proteome-wide genetic colocalisation tests, followed by single-instrument cis-MR analysis, and we report additional novel targets of therapeutic relevance. Finally, we examined associated phenotypes using the colocalising signals from the Open Targets Genetics portal (http://genetics.opentargets.org) to shed light on the biological basis of association of the proteins with the COVID-19 phenotypes. Materials and methods Key resources table Reagent type (species) or resourceDesignationSource or referenceIdentifiersAdditional informationCell line (Homo sapiens)HEK293-EYves Durocher, PMID:11788735RRID:CVCL_6974Transfected construct (Homo sapiens)pCMV6-CD209OrigeneCat.# SC304915Plasmid for CD209 cDNA expression in cell-based binding assayTransfected construct (Homo sapiens)pTT3-ACE2-BLHPMID:33432067Plasmid for recombinant ACE2 extracellular domain, for plate-based assays as the immobilised formTransfected construct (Homo sapiens)pTT3-CD209-BLHThis paperPlasmid for recombinant CD209 extracellular domain for plate-based assays as the immobilised formTransfected construct (Homo sapiens)pTT3-Cd4d3+ d4AddgeneRRID:Addgene_32402Plasmid for recombinant tag control (Cd4 domains 3 and 4)Transfected construct (Homo sapiens)pTT3-SPIKE-COMP-BLacThis paperPlasmid for recombinant SARS-CoV-2 spike extracellular domain for plate-based assays as the soluble formTransfected construct (Homo sapiens)pTT3-BirA-FLAGAddgeneRRID:Addgene_64395Biotin ligase plasmid for recombinant protein biotinylationPeptide, recombinant proteinStreptavidin R-phycoerythrinBioLegendCat.# 405245For tetramer staining in cell-based binding assayChemical compound, drugDAPI (4',6-diamidino-2-phenylindole)BioLegendCat.# 4228011 μM for flow cytometry live/dead stainingChemical compound, drugD-biotinSigma-AldrichCat.# 2031100 μM supplemented to cell culture media for biotinylationSoftware, algorithmR (version 4.0.3)R Foundationwww.r-project.orgRRID:SCR_001905Analysis and generating plots Genetic associations of proteins Request a detailed protocol We primarily used Sun et al. protein GWAS data (Sun et al., 2018; Emilsson et al., 2018) for the pan-/cis-MR analyses and for performing genetic colocalisation tests (described below). The pan-/cis-MR effects were expressed per standard deviation (SD) higher genetically predicted plasma protein concentrations. Two additional proteomic datasets (Emilsson et al., 2018; Suhre et al., 2017) were used to identify proteins associated with the ABO locus. The genotyping protocols and QC of these proteomic studies have been described previously (Sun et al., 2018; Emilsson et al., 2018; Suhre et al., 2017). All three of the proteomic studies have used the SOMAscan assay platform (an aptamer-based protein detection platform) to detect and quantify protein abundance (Gold et al., 2012). Genetic associations of COVID-19 Request a detailed protocol We used seven meta-analysed COVID-19 datasets from the October 2020 release of the ICDA COVID-HGI group (https://www.covid19hg.org/results/r4/). These seven COVID-19 outcomes are A1 (very severe respiratory confirmed COVID vs. not hospitalised COVID), A2 (very severe respiratory confirmed COVID vs. population), B1 (hospitalised COVID vs. not hospitalised COVID), B2 (hospitalised COVID vs. population), C1 (COVID vs. lab/self-reported negative), C2 (COVID vs. population), and D1 (predicted COVID from self-reported symptoms vs. predicted or self-reported non-COVID). Definitions of these outcomes are provided in Supplementary file 1. Harmonisation of protein and COVID summary statistics Request a detailed protocol Prior to analyses, we performed a liftover of datasets that reported genomic coordinates using the GRCh37 assembly to GRCh38. We also checked and ensured that the effect allele in a GWAS locus is the alternative allele in the forward strand of the reference genome. To infer strand for palindromic variants (variants with A/T or G/C alleles, i.e. variants with the same pair of letters on the forward strand as on the reverse strand), we first checked the orientation of all non-palindromic variants with respect to the reference genome to assess whether there was a strand consensus of 99% or more. For example, for a given GWAS, if ≥99% of the non-palindromic variants were on the forward strand, we assumed that the palindromic variant would also be on the forward strand; otherwise, they were excluded from analyses. Details of the harmonisation workflow are provided in our GitHub pages (EBISPOT, 2020; Opentargets Inc, 2021). Mendelian randomisation Request a detailed protocol To construct genetic instruments for MR analysis, we selected near-independent (r² = 0.05) genetic variants from across the genome ('pan'-instruments) or from within ±1 Mbp from the transcription start site (TSS) of the gene encoding the protein ('cis'-instruments) associated with the encoded protein abundance at p≤5 × 10–8 for pan-MR analyses and at a less stringent p ≤ 1 × 10⁻⁵ for cis-MR analyses (this p-value corrects for the number of proteins in the druggable genome Schmidt, 2020). We used the generalised summary data-based Mendelian randomisation (GSMR) approach with the heterogeneity-independent instrument (HEIDI)-outlier flag turned on to carry out the pan- and cis-MR analyses (Zhu et al., 2018). The GSMR software, using the HEIDI-outlier method, removes potentially pleiotropic instruments and accounts for the residual correlation between instruments (important as we are using near-independent genetic instruments). To select near-independent genetic instruments and account for linkage disequilibrium (LD) in the MR analyses, we used genotype data from 10,000 randomly sampled UK Biobank participants to create a reference LD matrix, which is ancestry-matched to the pQTL data we used. For each COVID-19 outcome, we used the Benjamini–Hochberg FDR (False Discovery Rate) threshold of 5% for significance, adjusting for 2042 tests in cis-MR analyses and 1286 tests in pan-MR analyses. For trans-acting instruments in pan-MR associations, variants were mapped to their respective cis-gene that had the highest overall V2G score in the Open Targets Genetics portal (Ghoussaini, 2021; Mountjoy, 2020; Open Targets Genetics, 2019a). Colocalisation analysis and phenome-wide association study Request a detailed protocol To identify shared causal genetic signals between protein and COVID outcomes, we used the Bayesian method of genetic colocalisation implemented in the coloc R package (Giambartolomei, 2014) using the marginal association statistics for each trait (i.e. assuming one independent signal in each region). We used beta and standard errors of cis-pQTLs of phenotype pairs as inputs. The default priors in coloc were used, that is, the prior of an SNP (single nucleotide polymorphism)-trait association is 1 × 10–4, and the prior of an SNP associating with both traits is 1 × 10–5. For each COVID-19 outcome, a posterior probability for shared causal genetic signal (PP.H4) threshold of more than 0.8 was used to identify shared causal genetic variants. For colocalising signals, we carried out a phenome-wide association study (PheWAS) using GWAS summary statistics (n = ~ 3000 GWAS) from the Open Targets Genetics portal (Ghoussaini, 2021; Mountjoy, 2020). Evidence against aptamer binding artefacts Request a detailed protocol For variants associated with proteins due to aptamer or epitope binding artefacts (which tend to be missense variants) (Joshi and Mayr, 2018), we first assessed whether genetic instruments for MR or coloc-based single-SNP MR analysis were associated with corresponding gene expression (i.e. whether they were also cis-eQTLs). This used gene expression data from the Open Targets Genetics portal (Ghoussaini, 2021). SNPs that were not cis-eQTLs were investigated further by identifying whether they were (or were in LD at r2 = 0.8 with) missense variants. To query if variants were missense or in LD with missense variants, we used the functional consequence data from Open Targets Genetics (Ghoussaini, 2021) (which used gnomAD v2 for variant effect prediction annotation, Lek, 2016). The reasoning was, if missense variants also had effects on corresponding gene expression, the causal inference using the missense variants as genetic instruments was unlikely to be biased even if the effect estimates were invalid. Where cis-pQTLs were not cis-eQTLs and were missense variants (or in LD with missense variants at r2 = 0.8) affecting the respective genes, these proteins were flagged and excluded from any further downstream analyses on the basis that the missense variant(s) might influence aptamer binding and produce biased effect estimates. Where cis-pQTLs were also cis-eQTLs and were missense variants (or in LD with missense variants) for the respective genes, although the effect estimates would not be valid, the causal inference using the instruments is unlikely to be biased; hence, these variants were retained in supplementary files and estimates of probes represented by these variants were flagged (using an asterisk) in the main figures. The rest, where cis-pQTLs had an effect on gene expression but were not missense variants or in LD with missense variants, were included in all analyses and presented without restrictions. Recombinant protein production Request a detailed protocol Recombinant human receptors and SARS-CoV-2 spike protein extracellular domains were expressed and purified as previously described (Shilts et al., 2021). Briefly, the full extracellular domain sequences of each were expressed as soluble secreted proteins in HEK293 cells. All proteins were affinity-purified using their hexahistidine tags. For biotinylated proteins, co-transfection of secreted BirA ligase in the presence of 100 µM D-biotin resulted in the covalent addition of a biotin group to an acceptor peptide tag, also as described previously (Kerr and Wright, 2012). The extracellular domain of CD209 (Q9NNX6) was defined as beginning at Pro114, while the full cDNA sequence was acquired from OriGene (#SC304915). Plate-based protein binding assay Request a detailed protocol The binding of biotinylated human receptor extracellular domains to pentameric SARS-CoV-2 spike protein was measured using the avidity-based extracellular interaction assay (AVEXIS) as previously described (Bushell et al., 2008). Briefly, the wells of a streptavidin-coated 96-well plate were saturated with biotinylated bait of either CD209, ACE2, or a previously described negative-control construct consisting only of the C-terminal protein tags shared by all other recombinant proteins (rat Cd4(d3 +4)-linker-Bio-6xHis) (Voulgaraki, 2005; Galaway and Wright, 2020). these we applied a of the full SARS-CoV-2 spike protein extracellular domain by a peptide sequence from the protein with a beta binding was measured by of a was by light at receptor binding assay Request a detailed protocol HEK293 cells were as described previously with expression encoding cDNA of CD209 or a the expression recombinant biotinylated spike protein was to as previously described et al., 2018). were with of spike or a control construct of protein tags on a flow as previously described (Shilts et al., 2021). HEK293 cell were provided by Research were first by All cell were for by and found to be all These cell are not by as availability Request a detailed protocol used to summary statistics are provided in (EBISPOT, 2020). for pan- and cis-MR analyses are provided on the GSMR for genetic colocalisation analyses are provided on the coloc GitHub 2021). All used in the to are provided in at 2021). Results and cis-MR analyses the of circulating ABO protein concentrations and soluble IL-6R in COVID-19 risk Our MR analysis used both genetic variants from across the genome and genetic variants near or in the gene encoding the relevant protein to associations of genetically predicted plasma protein concentrations with the risk of COVID-19 The COVID-19 are provided in Supplementary file 1. Although the pan-MR analysis leveraged genetic data from both and trans-acting a of from across the genome by HEIDI-outlier for some protein-COVID-19 pairs that were associated at 5% the associations with COVID-19 outcomes were driven by trans-acting or cis-acting genetic instruments. For example, although proteins were represented by both and trans-acting genetic and were represented only by cis-acting variants and one was driven entirely by trans-acting instruments ABO file the pan-MR analysis revealed protein probes associated with COVID outcomes at an FDR of 5% The selected by GSMR to represent these probes were also cis-eQTLs curated for the Open Targets Genetics portal Open Targets Genetics, and, the ABO signal be to as a that a nucleotide in the of were not missense variants or in LD with missense variants file the that SNPs with associations with proteins were used as genetic instruments for the of significant pan-MR 1 Open the of randomisation and cis-MR and genetic pan- and cis-MR methods used (Sun et al., 2018) as the of genetic instruments and the UK Biobank individual genotype data as reference We selected near-independent genetic instruments and performed MR analysis using generalised summary data-based Mendelian randomisation that for residual correlation between instruments. Genetic colocalisation analysis was used to posterior of shared causal genetic signal between protein and posterior probability of shared causal genetic signal of more than (i.e. a or posterior probability for 4 was used as evidence of genetic The line analysis the from target the three proteins with pan-MR evidence of association with COVID also had cis-MR evidence at cis-MR While the pan-MR analysis used genetic data from across the the cis-MR analysis genetic instrument to near 1 of or in the gene encoding the protein. proteins with pan-MR associations were supported by corresponding cis-MR associations and Supplementary file ABO, and these only ABO and IL-6R proteins had some evidence of genetic colocalisation with posterior (PP.H4) more than and of a shared genetic signal between protein and COVID-19 phenotype Although the of IL-6R was it had a = a signal of the IL-6R protein with the COVID-19 is a more likely than the association driven by independent Open associations of genetically predicted plasma protein concentrations with selected COVID-19 phenotypes. The estimates represent of COVID-19 per standard deviation (SD) of genetically predicted protein abundance using genetic instruments from across the genome randomisation The estimates represent of COVID per of genetically predicted protein abundance using genetic instruments near or in the gene encoding the protein represent The of the are to the of the of the For each COVID pan-MR associations at FDR 5% were a COVID phenotype a pQTL and the number of in the COVID phenotype the number of SNPs used as genetic instruments for the protein the posterior probability that protein and COVID traits the posterior probability evidence for vs. against shared causal variants and the candidate colocalising signal proteins that have that are either missense variants or in linkage disequilibrium with missense variants, their effect estimates potentially predicted ABO was associated with risk in out of seven COVID-19 outcomes These outcomes represented both susceptibility (e.g. COVID-19 vs. cis-MR per genetically predicted ABO × and severity (e.g. hospitalised COVID-19 vs. cis-MR p=1 × of COVID-19. predicted soluble IL-6R was only associated with higher risk of hospitalised COVID-19 compared to per genetically predicted × the SNPs involved in the pan-MR associations of the all probes IL-6R and ABO had at one trans-acting and in all these at one of the trans-acting SNPs were to the ABO gene by the Open Targets Genetics V2G the of the ABO genetic Furthermore, when the of pan-MR associations of these probes across all seven COVID-19 outcomes, the protein probes that have trans-acting ABO SNPs a similar association as the ABO protein associated with only COVID-19 outcomes that have file 1 of proteins reported in our study and the of evidence their by by by and single-SNP of cis-acting of trans-acting

  • Discussion
  • Cite Count Icon 16
  • 10.1016/j.jhep.2022.05.005
Associations of muscle mass and grip strength with severe NAFLD: A prospective study of 333,295 UK Biobank participants.
  • Nov 1, 2022
  • Journal of Hepatology
  • Lanlan Chen + 2 more

Associations of muscle mass and grip strength with severe NAFLD: A prospective study of 333,295 UK Biobank participants.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 1
  • 10.1155/2022/4120711
Analysis of Trends in Awareness Regarding Hepatitis Using Bayesian Multiple Logistic Regression Model
  • Jun 9, 2022
  • Mathematical Problems in Engineering
  • Ali Al-Alwan + 4 more

Diseases like hepatitis remained a major health concern, especially in developing countries. The awareness and knowledge about such diseases are of prime importance. The analysis of socioeconomic factors associated with the tendency of awareness and knowledge about the said diseases is fundamental. However, in developing countries like Pakistan, very few studies have considered such investigations using nationally representative data. In addition, a careful review of the literature suggests that no studies have analyzed the trends in awareness and knowledge about said disease with respect to time using nationally representative datasets. Furthermore, the existing literature regarding these studies has utilized the classical methods for the analysis. We have considered a detailed study for analyzing the trends in awareness and knowledge about the said disease in the general population of Pakistan from 2012 to 2018, using nationally representative data collected through Pakistan Demographic and Health Surveys. In addition, we have considered the Bayesian methods for the analysis and performance of the proposed Bayes methods that have been compared with the frequently used classical methods. The results indicated that the proposed Bayesian multiple logistic regression models performed better as compared to classical multiple logistic regression models (CMLRMs). This is due to fact that widths of 95% CIs were smaller for Bayesian multiple logistic regression models (BMLRM), as compared to classical multiple logistic regression models. The findings of the study suggest that there are severe disparities (with respect to different socioeconomic groups) in the knowledge and awareness of respondents for hepatitis. The levels of knowledge and awareness about the said disease are drastically low for respondents living in rural areas, having lower levels of education and wealth. These disparities seem to persist, as the corresponding odds have not changed much during the period 2012 to 2018. The policy-maker should plan and implement the strategies to reduce the observed disparities for different sectors of society.

  • Research Article
  • Cite Count Icon 19
  • 10.1016/j.jse.2020.11.025
A genome-wide association study for shoulder impingement and rotator cuff disease.
  • Jan 19, 2021
  • Journal of Shoulder and Elbow Surgery
  • Stuart K Kim + 3 more

A genome-wide association study for shoulder impingement and rotator cuff disease.

  • Research Article
  • Cite Count Icon 14
  • 10.1161/hcg.0000000000000046
Interdisciplinary Models for Research and Clinical Endeavors in Genomic Medicine: A Scientific Statement From the American Heart Association.
  • Jun 1, 2018
  • Circulation: Genomic and Precision Medicine
  • Kiran Musunuru + 14 more

The completion of the Human Genome Project has unleashed a wealth of human genomics information, but it remains unclear how best to implement this information for the benefit of patients. The standard approach of biomedical research, with researchers pursuing advances in knowledge in the laboratory and, separately, clinicians translating research findings into the clinic as much as decades later, will need to give way to new interdisciplinary models for research in genomic medicine. These models should include scientists and clinicians actively working as teams to study patients and populations recruited in clinical settings and communities to make genomics discoveries-through the combined efforts of data scientists, clinical researchers, epidemiologists, and basic scientists-and to rapidly apply these discoveries in the clinic for the prediction, prevention, diagnosis, prognosis, and treatment of cardiovascular diseases and stroke. The highly publicized US Precision Medicine Initiative, also known as All of Us, is a large-scale program funded by the US National Institutes of Health that will energize these efforts, but several ongoing studies such as the UK Biobank Initiative; the Million Veteran Program; the Electronic Medical Records and Genomics Network; the Kaiser Permanente Research Program on Genes, Environment and Health; and the DiscovEHR collaboration are already providing exemplary models of this kind of interdisciplinary work. In this statement, we outline the opportunities and challenges in broadly implementing new interdisciplinary models in academic medical centers and community settings and bringing the promise of genomics to fruition.

  • Research Article
  • Cite Count Icon 2
  • 10.1002/agj2.21670
Soybean prediction using computationally efficient Bayesian spatial regression models and satellite imagery
  • Sep 3, 2024
  • Agronomy Journal
  • Richard J Fischer + 4 more

Preharvest yield estimates can be used for harvest planning, marketing, and prescribing in‐season fertilizer and pesticide applications. One approach that is being widely tested is the use of machine learning (ML) or artificial intelligence (AI) algorithms to estimate yields. However, one barrier to the adoption of this approach is that ML/AI algorithms behave as a black block. An alternative approach is to create an algorithm using Bayesian statistics. In Bayesian statistics, prior information is used to help create the algorithm. However, algorithms based on Bayesian statistics are not often computationally efficient. The objective of the current study was to compare the accuracy and computational efficiency of four Bayesian models that used different assumptions to reduce the execution time. In this paper, the Bayesian multiple linear regression (BLR), Bayesian spatial, Bayesian skewed spatial regression, and the Bayesian nearest neighbor Gaussian process (NNGP) models were compared with ML non‐Bayesian random forest model. In this analysis, soybean (Glycine max) yields were the response variable (y), and spaced‐based blue, green, red, and near‐infrared reflectance that was measured with the PlanetScope satellite were the predictor (x). Among the models tested, the Bayesian (NNGP; R2‐testing = 0.485) model, which captures the short‐range correlation, outperformed the (BLR; R2‐testing = 0.02), Bayesian spatial regression (SRM; R2‐testing = 0.087), and Bayesian skewed spatial regression (sSRM; R2‐testing = 0.236) models. However, associated with improved accuracy was an increase in run time from 534 s for the BLR model to 2047 s for the NNGP model. These data show that relatively accurate within‐field yield estimates can be obtained without sacrificing computational efficiency and that the coefficients have biological meaning. However, all Bayesian models had lower R2 values and higher execution times than the random forest model.

  • Research Article
  • Cite Count Icon 10
  • 10.3741/jkwra.2008.41.3.325
Bayesian 다중회귀분석을 이용한 저수량(Low flow) 지역 빈도분석
  • Mar 15, 2008
  • Journal of Korea Water Resources Association
  • Sang-Ug Kim + 1 more

본 연구는 저수량 지역 빈도분석(regional low flow frequency analysis)을 수행하기 위하여 일반최소자승법(ordinary least squares method)을 이용한 Bayesian 다중회귀분석을 적용하였으며, 불확실성측면에서의 효과를 탐색하기 위하여 Bayesian 다중회귀분석에 의한 추정치와 t 분포를 이용하여 산정한 일반 다중회귀분석의 추정치의 신뢰구간을 비교분석하였다. 각 재현기간별 비교결과를 보면 t 분포를 이용하여 산정된 평균 추정치와 Bayesian 다중회귀분석에 의한 평균 추정치는 크게 다르지 않았다. 그러나 불확실성 측면에서 평가해볼 때 신뢰구간의 상한추정치와 하한추정치의 차이는 Bayesian 다중회귀분석을 사용한 경우가 기존 방법을 사용한 경우보다 훨씬 작은 것으로 나타났으며, 이로부터 저수량(low flow) 지역 빈도분석을 수행하는 경우 Bayesian 다중회귀분석이 일반 회귀분석보다 불확실성을 표현하는데 있어서 우수하다는 결과를 얻을 수 있었다. 또한 낙동강 유역에 2개의 미계측 유역을 선정하고 구축된 Bayesian 다중회귀모형을 적용하여 불확실성을 포함한 미계측 유역에서의 저수량(low flow)을 추정하였으며 이와 같은 방법이 미계측 유역에서의 저수(low flow) 특성을 나타내는 데 있어서 효과적일 수 있음을 입증하였다. This study employs Bayesian multiple regression analysis using the ordinary least squares method for regional low flow frequency analysis. The parameter estimates using the Bayesian multiple regression analysis were compared to conventional analysis using the t-distribution. In these comparisons, the mean values from the t-distribution and the Bayesian analysis at each return period are not significantly different. However, the difference between upper and lower limits is remarkably reduced using the Bayesian multiple regression. Therefore, from the point of view of uncertainty analysis, Bayesian multiple regression analysis is more attractive than the conventional method based on a t-distribution because the low flow sample size at the site of interest is typically insufficient to perform low flow frequency analysis. Also, we performed low flow prediction, including confidence interval, at two ungauged catchments in the Nakdong River basin using the developed Bayesian multiple regression model. The Bayesian prediction proves effective to infer the low flow characteristic at the ungauged catchment.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 4
  • 10.1186/s13104-023-06438-4
Deriving GWAS summary estimates for paternal smoking in UK biobank: a GWAS by subtraction
  • Jul 30, 2023
  • BMC Research Notes
  • Benjamin Woolf + 3 more

ObjectiveTo use genome-wide association study (GWAS) by subtraction, a method for deriving novel GWASs from existing summary statistics, to derive genome-wide summary statistics for paternal smoking.ResultA GWAS by subtraction was implemented using a weighted linear model that defined the child-genotype paternal-phenotype association as the child-genotype child-phenotype association minus the child-genotype maternal-phenotype association. We first use the laws of inherence to derive the weighted linear model. We then implemented the linear model to create a GWAS of paternal smoking by subtracting the summary statistics from a GWAS of maternal smoking from the summary statistics of a GWAS of the index individual’s smoking. We used a Monte-Carlo simulation to validate the model and showed that this approach performed similarly in terms of bias to performing a traditional GWAS of paternal smoking. Finally, we validated the summary statistics in a Mendelian randomisation analysis by demonstrating an association of genetically predicted paternal smoking with paternal lung cancer and emphysema.

  • Research Article
  • 10.1093/ije/dyab168.180
649Personal history of keratinocyte carcinoma is a marker of inherited cancer risk: Mendelian randomization analyses
  • Sep 1, 2021
  • International Journal of Epidemiology
  • Jean Claude Dusingize + 9 more

Background A personal history of keratinocyte carcinoma (KC) has been reported as a risk factor for developing subsequent primary cutaneous and non-cutaneous malignancies. However most evidence to date stems from observational studies which are prone to bias, confounding and reverse causation. Our aim was to examine this association using different Mendelian randomization (MR) approaches. Methods We performed a one-sample MR analysis using individual-level data from the UK Biobank (n = 394,306). This analysis was then validated in an independent dataset in the QSkin cohort (n = 16,896). Using 64 independent genetic variants known to be associated with KC, we generated a polygenic risk score (PRS) for each participant in the UK Biobank and the QSkin cohort. We then performed two-sample MR analyses using genome-wide association study (GWAS) summary statistics. We tested the association between genetically predicted KC and risk of subsequent cancer using logistic regression. Results Results from one-sample MR analyses in the UK Biobank indicated that a personal history of KC was significantly associated with cancer overall (excluding melanoma) (OR: 1.15, 95% CI: 1.10-1.20, per doubling the prevalence of KC). The results from the two-sample MR corroborate the findings from the one-sample MR, although the risk estimate was lower (OR: 1.05, 95% CI: 1.03-1.07). Conclusions Our MR analyses suggest that genetically predicted KC is a risk factor for developing subsequent primary malignancies. Key messages A personal history of KC may serve as a proxy marker of inherited cancer risk.

  • Research Article
  • Cite Count Icon 11
  • 10.1016/j.jse.2022.07.005
A shared genetic architecture between adhesive capsulitis and Dupuytren disease.
  • Jan 1, 2023
  • Journal of Shoulder and Elbow Surgery
  • Stuart K Kim + 3 more

A shared genetic architecture between adhesive capsulitis and Dupuytren disease.

  • Peer Review Report
  • 10.7554/elife.79348.sa1
Decision letter: Integrative analysis of metabolite GWAS illuminates the molecular basis of pleiotropy and genetic correlation
  • Jun 18, 2022
  • Alexander Young + 1 more

Article Figures and data Abstract Editor's evaluation Introduction Results Discussion Methods Data availability References Decision letter Author response Article and author information Metrics Abstract Pleiotropy and genetic correlation are widespread features in genome-wide association studies (GWAS), but they are often difficult to interpret at the molecular level. Here, we perform GWAS of 16 metabolites clustered at the intersection of amino acid catabolism, glycolysis, and ketone body metabolism in a subset of UK Biobank. We utilize the well-documented biochemistry jointly impacting these metabolites to analyze pleiotropic effects in the context of their pathways. Among the 213 lead GWAS hits, we find a strong enrichment for genes encoding pathway-relevant enzymes and transporters. We demonstrate that the effect directions of variants acting on biology between metabolite pairs often contrast with those of upstream or downstream variants as well as the polygenic background. Thus, we find that these outlier variants often reflect biology local to the traits. Finally, we explore the implications for interpreting disease GWAS, underscoring the potential of unifying biochemistry with dense metabolomics data to understand the molecular basis of pleiotropy in complex traits and diseases. Editor's evaluation Smith and colleagues provide a framework for understanding a seemingly paradoxical observation in human genetics: two phenotypes may be closely correlated to each other, and the patterns of genetic variation that influence both phenotypes may be widely shared at the genome-wide level, but there are often specific genetic variants that show discordant patterns. Though the observations in this article are derived from analysis of metabolic phenotypes, this may have broader relevance to interpreting the results from disease-related genetic association studies, and shed light on the processes that connect different disease phenotypes. https://doi.org/10.7554/eLife.79348.sa0 Decision letter Reviews on Sciety eLife's review process Introduction A central challenge in the field of human genetics is understanding the mechanism of how genetic variants influence complex traits and diseases. Genome-wide association studies (GWAS) have begun characterizing the genetic architecture of complex traits, but the molecular mechanisms connecting genetic variants to these traits are rarely understood. This is particularly true for understanding pleiotropy, when a variant affects multiple traits (Solovieff et al., 2013). It is possible to estimate the genetic correlation between traits (Bulik-Sullivan et al., 2015b; Shi et al., 2017), but it is often unclear what contributes to this at a molecular or physiological level. A handful of in vitro disease-focused 'post-GWAS' studies have convincingly shown the mechanisms driving pleiotropy of individual key associations (Warren et al., 2017; Sinnott-Armstrong et al., 2021b); however, these studies are highly specific and time-consuming. Developing statistical and computational approaches to identify putative molecular mechanisms is invaluable to advancing our understanding of where and how pleiotropic GWAS variants act. In this study, we use metabolites as model traits to understand pleiotropic features of genetic architecture. Metabolites are small molecules interconverted by a series of biochemical pathways and are an appealing model system for studying pleiotropy because their pathways are typically well-documented and biologically simpler than those underlying other complex traits (Gieger et al., 2008; Sinnott-Armstrong et al., 2021a). Previous work in Mendelian genetics has identified inborn errors of metabolism (IEM) in many enzymes (Woidy et al., 2018). Metabolite GWAS, which have long observed pervasive pleiotropy at these IEM genes and other loci (Shin et al., 2014; Yeung, 2021), offer a potential opportunity to further explore the relationships between intermediate molecules and disease outcomes at scale. Here, we jointly analyzed GWAS results of 16 plasma metabolites from the Nightingale Health Nuclear Magnetic Resonance (NMR) Spectroscopy platform in nearly 100,000 individuals in the UK Biobank (Julkunen et al., 2021; Figure 1; see 'Methods'). These 16 metabolites included glucose, pyruvate, lactate, citrate, isoleucine, leucine, valine, alanine, phenylalanine, tyrosine, glutamine, histidine, glycine, acetoacetate, acetone, and 3-hydroxybutyrate. They were chosen based on their biochemical proximity to each other, their relevance to health and disease, and because the genes and enzymes involved in their metabolism are well-characterized. They play especially important roles in energy generation and energy storage pathways such as glycolysis, the citric acid cycle, amino acid metabolism, and ketone body formation. They are relevant to many metabolic diseases including type 2 diabetes (Laffel, 1999; Guasch-Ferré et al., 2020; Newsholme et al., 2007), cardiovascular disease (Lusis and Weiss, 2010), and non-alcoholic fatty liver disease (Watt et al., 2019). Figure 1 Download asset Open asset Biochemistry of relevant metabolites. Pathway diagram and molecular structure of relevant metabolites, colored by their biochemical groups. The pathway diagram was curated from multiple resources (see 'Methods'). All solid lines represent a single chemical reaction step. Dotted lines represent a simplification of multiple steps. For simplicity, only a subset of all the reactions each metabolite participates in is shown. Genes encoding the enzymes that catalyze the above chemical reactions are known and presented in Figure 3—figure supplement 1. Numerous GWAS have begun characterizing the genetic architecture of metabolites and found them to be heritable and polygenic (Lemaitre et al., 2011; Suhre et al., 2011). Recent metabolite studies have shown that leveraging information about the biochemical pathways relevant to a given metabolite (Teslovich et al., 2018; Graham et al., 2021; Rueedi et al., 2017; Sinnott-Armstrong et al., 2021a) can allow for more interpretable gene annotation of GWAS hits. This has led to the dissection of individual associations of biomarkers, such as lipids (Kettunen et al., 2016), glycine (Wittemans et al., 2019), and intermediate clinical measures (Lotta et al., 2021), with cardiometabolic and other diseases. The pervasive pleiotropy at these GWAS loci with other metabolites as well as disease (Lotta et al., 2021; Pott et al., 2019) suggests the potential of utilizing these data for investigating the mechanism of pleiotropic effects as a core component of genetic architecture. While recent GWAS have begun jointly investigating multiple metabolites (Cichonska et al., 2016; Ruotsalainen et al., 2021; Qi and Chatterjee, 2018), they have yet to do so in the context of their biochemical pathways. In this article, we demonstrate that investigating the effects of pleiotropic variants on biologically related metabolites allows for a better understanding of why these variants have their observed joint effects. Our results reveal striking heterogeneity in genetic correlation across the genome and provide a biologically intuitive basis for understanding this heterogeneity. Together, this allows us to dissect the molecular basis of metabolic disease GWAS variants and enables us to directly define the mechanism relating an example variant to its associated disease. Results Insights into the shared genetic architecture of biologically related metabolites We chose 16 metabolites from the 249 available through the Nightingale NMR platform in a subset of the UK Biobank (Figure 1; see 'Methods'). These 16 metabolites were selected based on their biochemical proximity, relevance to health and disease, and because the genes and enzymes involved in their metabolism are well-characterized. We classified the 16 metabolites into four groups based on shared biochemistry: Glycolysis (glucose, pyruvate, lactate, citrate), Branched Chain Amino Acid (BCAA; isoleucine, leucine, valine), Other Amino Acid (alanine, phenylalanine, tyrosine, glutamine, histidine, glycine), and Ketone Body (acetoacetate, acetone, 3-hydroxybutyrate). Trait measurements were log-transformed and adjusted for relevant technical covariates. After outlier removal, we obtained a primary dataset of 94,464 genotyped European-ancestry individuals with data for all 16 metabolites. We first sought to characterize the genetic architecture underlying these metabolites by performing GWAS for each (Figure 2—figure supplement 1). Hits from individual GWAS were clumped with an r2 of 0.01 per megabase, combined across metabolites, then pruned to the single nucleotide polymorphism (SNP) with the most significant p-value within 0.1 cM. This resulted in 213 lead variants with a genome-wide significant association in at least one metabolite, referred to as the metabolite GWAS hits. Glycine had the largest number of significant associations with 77 hits (Figure 2a). There were 47 variants with significant associations in more than one metabolite, including rs2939302 (near the gene GLS2), which was significant in 9 of the 16 metabolites, and rs1260326 (GCKR), which was significant in 8. Glycine also had the highest Heritability Estimate from Summary Statistics (HESS) total SNP heritability of 0.284 (Supplementary file 1 and Figure 2—figure supplement 2). Figure 2 with 4 supplements see all Download asset Open asset Overview genetic architecture of metabolites. (a) Number of genome-wide association studies (GWAS) hits per metabolite. (b) Pairwise LDSC genetic correlations between the metabolites, clustered by genetic correlation. (c) Mendelian randomization weighted results between the metabolites. (d) Biclustered standardized effect size in each metabolite for the 213 metabolite GWAS hits. For visualization, effect sizes were divided by the standard error then inverse normal transformed and standardized. Each variant was aligned to have a positive median score across metabolites. To understand the shared genetics of these metabolites, we then investigated the extent of pleiotropy between and within biochemical groups. In order to examine this, we first calculated pairwise LD Score regression (LDSC) genetic correlation across the 16 metabolites. We found substantial genome-wide sharing for many pairs of metabolites, especially for metabolites within the same biochemical group (Figure 2b; phenotypic correlation in Figure 2—figure supplement 3). We then explored pleiotropic effects beyond the polygenic background by examining the structure within the metabolite GWAS hits. Pairwise Mendelian randomization (MR) between the metabolites emphasized the intertwined nature of these traits (Figure 2c). Despite only taking into account genetic effects, MR largely clustered metabolites in a way that reflects their biochemical groups. The extensive pleiotropy across the 16 traits, with similar sharing inside biochemical groups, is also illustrated by the structure visible in the normalized effect sizes for each metabolite GWAS hit (Figure 2d). Together, these analyses support substantial, but not always consistent, genetic overlap between the traits, particularly in the polygenic components. In the remainder of this article, we will seek a deeper understanding of the biochemical relationships between genotypes and metabolite levels. Characterizing the biological functions of candidate genes An important step in understanding the pathway-level mechanisms of variants is knowing which gene a variant is affecting and how that gene relates to the biology of the pathway. Different types of genes influence trait biology through distinct mechanisms. Metabolite biology is documented in genetic and biochemical databases based on the extensive history of biochemical research (Supplementary file 2). Thus, we developed a pipeline for annotating the 213 metabolite GWAS hits with a single most likely gene using gene proximity and manual curation of these databases (Supplementary file 3 and Figure 2—figure supplement 4; see 'Methods'). We annotated 68 variants with genes encoding pathway-relevant enzymes (25-fold enrichment, 95% CI [20-fold, 33-fold], Poisson rate test p<2e-16), 46 with genes encoding transporters (5.2-fold enrichment, 95% CI [3.7-fold, 7.2-fold], p=9e-16), and 30 with genes encoding transcription factors (TFs; 7-fold enrichment among liver marker TFs, 95% CI [3.0-fold, 14-fold], p=3e-5; Figure 3). Overall, 69% of variants were assigned to the closest gene and 49% of variants assigned to a pathway-relevant enzyme gene were assigned known IEM genes (Woidy et al., 2018). The substantial enrichment for biologically interpretable variants suggests that examining the genetic basis of these traits will allow for the development of hypotheses around relevant molecular mechanisms underlying pleiotropy. Figure 3 with 2 supplements see all Download asset Open asset Gene annotation of metabolite genome-wide association studies (GWAS) hits. Each gene is colored based on the biochemical group with the most associated metabolites (p<1e-4). If multiple biochemical groups are tied for the most associations for a given gene, they are all shown. (a) Expanded pathway diagram with all genes (italicized) that encode pathway-relevant enzymes and were a metabolite GWAS hits. (b) List of all genes of the metabolite GWAS hits that encode transporters and transcription factors (TFs). There were 69 metabolite GWAS hits that are not shown. Of these, 60 were annotated with genes assigned to the gene type general cell function (14 of these 60 were related to lipid function), and 9 were assigned to a gene of unknown function or that did not have any genes nearby (see 'Methods'). Next we sought to understand which genes and subpathways were most relevant to each biochemical group. We assigned each gene to the biochemical group with the most associated metabolites (Supplementary file 4; see 'Methods'). Genes were largely assigned to the group whose relevant biology was nearest the protein encoded by the gene. For example, BCAT2 encodes an enzyme responsible for the first step in the breakdown of all three BCAAs and was assigned to the BCAA group. OXCT1 encodes an enzyme responsible for the conversion of acetoacetyl-CoA to the ketone body acetoacetate and was assigned to the Ketone Body group. Similarly, SLC7A9 encodes a protein that transports amino acids and was assigned to the Other Amino Acid group, while TCF7L2 is a TF assigned to the Glycolysis group and involved in blood glucose homeostasis. These results confirm that these variants are affecting known trait-relevant biology and reflecting the local structure of these pathways. Interestingly, a large fraction of the genes involved in trait-relevant biology were genome-wide significant hits for at least one of the 16 metabolites. Specifically, of the 139 total genes encoding enzymes in the pathway diagram for these metabolites (Figure 3—figure supplement 1), 51 genes had at least one GWAS hit. Additionally, we performed an ancestry-inclusive GWAS of all 98,189 individuals with complete metabolite data for follow-up analysis. In this ancestry-inclusive analysis, we identified 41 additional hits not found in the European-only GWAS, including associations at 7 additional pathway-relevant genes (Supplementary file 5 and Figure 3—figure supplement 2). This highlights the potential for large-scale, ancestry-inclusive GWAS to discover more biochemically relevant associations among these traits. Together, these findings suggest that GWAS reflect, and have the potential to illuminate, the complex biochemical pathways interconverting these metabolites. Investigating the mechanisms of pleiotropy in trait pairs Given the overlap between the biology of these metabolites and their hits, we next sought to understand the molecular causes of pleiotropy in trait pairs. We found 26 genetically correlated metabolite pairs at a local false sign rate < 0.005. For example, alanine and its strongest genetic correlation partner, isoleucine, share a genetic correlation of rg = 0.52 (SE = 0.05, p=9e-23). Similarly, plotting the effects of the 213 GWAS variants on these two traits indicates a strong positive correlation (Figure 4a). Nonetheless, we noted several outlier loci, including rs370014171 (PDPR) and rs77010315 (SLC36A2), which have strong discordant effects. We were intrigued to understand why these two variants had discordant effects on alanine and isoleucine relative to their overall positive genetic correlation, while the majority of other variants had concordant effects. Figure 4 with 7 supplements see all Download asset Open asset Discordant variant analysis. (a) Comparison between the effects on alanine levels versus the effects on isoleucine levels for the 213 metabolite genome-wide association studies (GWAS) hits. This highlights two discordant variants: rs370014171 (PDPR) and rs77010315 (SLC36A2). Variants without an association p<1e-4 in both metabolites are labeled 'Neither.' (b) Graphical representation of where these two discordant variants act in the pathway, represented by green stars, relative to other upstream variants driving the positive genetic correlation. Below are the hypothesized mechanisms explaining the GWAS results in relevant metabolites for each of these discordant variants. Data are shown in circles with the coloring corresponding to the effect (beta) of that variant on that metabolite. A black outline represents an association with p<1e-4. Orange text and arrows represent a hypothesized increase (direction, not magnitude) in flux and blue corresponds to a decrease. (c) Results for rs370014171 near the gene PDPR, which encodes a protein that activates the conversion of pyruvate to acetyl-CoA. All solid lines represent a single chemical reaction step. Dotted lines represent a simplification of multiple steps. (d) Results for rs77010315 in the gene SLC36A2, which encodes a small amino acid transporter. Outlier variants are appealing case studies for understanding the molecular basis of pleiotropy because they affect traits in an exceptional way. Thus, we reasoned that understanding large-effect variants inconsistent with the global genetic correlation would reflect interesting biology relevant to the traits. For example, the proteins encoded by PDPR and SLC36A2 are both located between alanine and isoleucine in the biochemical pathway (Figure 4b). This suggests that where variants act in the pathway may influence the direction of effect they have on metabolites. To better understand how these two variants affect alanine and isoleucine and explain their outlier behavior, we examined their effect size and direction in the context of their location in the pathway. We then used the variants' metabolite associations to develop candidate mechanisms for how each variant could be jointly influencing the levels of these metabolites. As an illustration, we first consider variant rs370014171. This variant was assigned to gene PDPR because it was the second closest gene, the closest pathway-relevant enzyme, and within 100 kb (12.3 kb to its gene boundaries). PDPR activates the enzyme that catalyzes the conversion of pyruvate to acetyl-CoA (Figure 4c, Figure 4—figure supplement 1). A candidate mechanism for this variant, supported by the effect size and direction for the 16 metabolites where relevant, is that it increases PDPR activity. There was colocalization of association signals across the five significant metabolites using both conditional SNP-level analyses (Figure 4—figure supplement 2) and running coloc once adjusting for secondary signals at alanine ( Figure 4—figure supplement 3, Figure 4—figure supplement 4). This would lead to increased conversion of pyruvate to acetyl-CoA and thus decreased pyruvate (β = -0.023 SDs, SE = 0.003, p=3e-20). To compensate for the subsequent decreased pyruvate levels, there would be increased conversion of alanine to pyruvate causing a decrease in alanine. In response to the increased acetyl-CoA, there would be decreased breakdown of metabolites normally catabolized for its production, including isoleucine, resulting in an increase in isoleucine levels. Thus, this variant has an opposite effect on alanine and isoleucine, despite their overall positive genetic correlation, likely because it affects the activity of an enzyme that acts in the pathway between the pair of metabolites. As expected due to the high correlation between the levels of the three BCAAs, this variant is also a discordant variant for alanine with valine (rg = 0.51, SE = 0.05, p=2e-21), and alanine with leucine (rg = 0.49, SE = 0.06, p=1e-16). As a second example, variant rs77010315 is a missense variant in SLC36A2. SLC36A2 encodes a transporter for small amino acids such as alanine (Figure 4d, Figure 4—figure supplement 5). There was colocalization of association signals across the four significant metabolites using both conditional SNP-level analyses (Figure 4—figure supplement 6) and running coloc (Figure 4—figure supplement 7). A candidate mechanism explaining the observed metabolite associations in our data and outlier behavior for this variant is that it increases transport of alanine into cells by SLC36A2. This would result in a decrease in levels of alanine in the blood, but an increase of alanine in cells. This additional intracellular alanine would then allow for increased conversion of alanine to pyruvate, thereby increasing levels of downstream metabolites in the blood, including isoleucine. Thus, this variant has an opposite effect on alanine and isoleucine, despite their overall positive genetic correlation, but in this case because it affects biology between the metabolites at the transporter level. Quantifying global properties of molecular pleiotropy Based on these results, we hypothesized that the two variants described above, and others like them, exhibit outlier behavior because they affect biology between the two metabolites (Figure 5a). We consider biology 'between' a given pair of metabolites as the biochemical connecting them, which can a one metabolite to the other, as well as other such as those This is because all biochemical including can have relationships due to (see for Figure supplement 1). correlation reflects the direction of effect that most associated variants have on two traits. when two metabolites are biologically near each other, the 'between' biology is such that only a of variants directly affect the 'between' Figure 5 with 2 supplements see all Download asset Open asset of discordant and concordant variants. (a) model for the mechanism of a discordant This example is for a discordant variant that has opposite effect directions on a pair of metabolites with a positive overall genetic correlation because it affects biology between (b) model for the mechanism of a concordant This example is for a concordant variant that has the same effect direction on a pair of metabolites with a positive overall genetic correlation because it affects biology upstream both metabolites. (c) of the discordant and concordant variants that have a pathway-relevant enzyme or transporter annotation versus those with a different annotation total = Discordant variants are for the gene types of pathway-relevant enzyme or as would be expected in the model of discordant variants affecting biology between metabolites. (d) of the discordant and concordant variants annotated with a pathway-relevant enzyme that affect biology between versus not between their significant metabolite pairs total = were performed using and the are from 95% CI calculated by score Thus, we hypothesized that the genetic correlation of two biologically related metabolites reflects the effects of variants upstream or downstream of the metabolites, the effects of those We developed an that variants affecting biology upstream or downstream of the two metabolites have concordant effects (Figure While the overall genetic correlation for two biologically related metabolites can also be due to factors such as In this variants acting between the two metabolites would have the same direction of effect on both metabolites, them discordant with the overall genetic correlation (Figure supplement 2). To these we based on the of their effects with the overall LDSC genetic correlation. If a variant had an effect direction opposite the overall LDSC genetic correlation in at least one significant metabolite pair in p<1e-4 in the it was classified as For example, a discordant variant for a metabolite pair with a positive genetic correlation would have a association in one of the metabolites and a positive association in the If a variant had an effect direction with the overall genetic correlation for its significant metabolite it was classified as Variants without multiple or where associated traits were not genetically were classified as In of the metabolite GWAS hits that had at least one significant metabolite we found 26 total discordant pairs across variants (Supplementary file We then investigated overall properties of discordant variants relative to concordant We that discordant variants are more likely to affect genes encoding enzymes and transporters than all other genes including TFs, general cell function and those of unknown function = 95% CI Figure This is in contrast to concordant which do not show an enrichment for enzymes and transporters relative to other gene types = 95% CI These observations are with our model that discordant variants to affect biology between relevant pairs of metabolites and general cell function genes act these metabolic pathways. Thus, they are more likely to affect biology upstream or downstream of both metabolites. In for variants affecting pathway-relevant where the location in the pathway that the variant is acting relative to the metabolites is we were to directly test our We found that discordant variants affecting pathway-relevant enzymes are more likely to act than upstream or the metabolites for which they are discordant = 95% CI Figure We then sought to this by a model the effects of all variants affecting between versus biology at a pathway and genome-wide (Figure In this model that pathways biology between two metabolites will have a local genetic correlation opposite that of nearby pathways and that the of both will that of the global polygenic background. As a case study, we on alanine and glutamine, which have a positive overall genetic correlation (rg = SE = Figure Figure supplement 1). We then et al., on variants within 100 kb of genes in each pathway and the corresponding local genetic correlations (see 'Methods'). Figure with 1 supplement see all Download asset Open asset genetic correlation. (a) of expected local genetic correlation with effects of variants affecting 'between' versus biology at a pathway and genome-wide level. (b) LD Score the polygenic correlation of alanine and For the LD were into The the and SE within each genome and overall genetic correlation was calculated using the standard error = (c) Results for the local genetic correlation of alanine and for variants within 100 kb of genes in each pathway. errors are shown = the only enzymes and are pathway-relevant enzymes for these metabolites. Summary for and other can be found in file (d) Pathway diagram the pathways included in the local genetic correlation analysis and the of their genes relative to alanine and body genes were from (c) because the number of genes they to All arrows and in the are and shown for We found that the local genetic correlations around genes in the and Acid Pathway and around genes in the Other Amino Acid Pathway were (Figure of these pathways genes affecting biology between alanine and (Figure In striking nearby such as the had a positive local genetic

  • Research Article
  • Cite Count Icon 11
  • 10.1016/j.csda.2012.02.019
Bayesian multiple response kernel regression model for high dimensional data and its practical applications in near infrared spectroscopy
  • Mar 3, 2012
  • Computational Statistics &amp; Data Analysis
  • Sounak Chakraborty

Bayesian multiple response kernel regression model for high dimensional data and its practical applications in near infrared spectroscopy

  • Discussion
  • Cite Count Icon 4
  • 10.1016/j.jhep.2022.10.032
Assessing causal relationship between non-alcoholic fatty liver disease and risk of atrial fibrillation
  • Nov 10, 2022
  • Journal of Hepatology
  • Ziang Li + 4 more

Assessing causal relationship between non-alcoholic fatty liver disease and risk of atrial fibrillation

Save Icon
Up Arrow
Open/Close
Setting-up Chat
Loading Interface