Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

PRSice-2: Polygenic Risk Score software for biobank-scale data.

  • Abstract
  • Highlights & Summary
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

BackgroundPolygenic risk score (PRS) analyses have become an integral part of biomedical research, exploited to gain insights into shared aetiology among traits, to control for genomic profile in experimental studies, and to strengthen causal inference, among a range of applications. Substantial efforts are now devoted to biobank projects to collect large genetic and phenotypic data, providing unprecedented opportunity for genetic discovery and applications. To process the large-scale data provided by such biobank resources, highly efficient and scalable methods and software are required.ResultsHere we introduce PRSice-2, an efficient and scalable software program for automating and simplifying PRS analyses on large-scale data. PRSice-2 handles both genotyped and imputed data, provides empirical association P-values free from inflation due to overfitting, supports different inheritance models, and can evaluate multiple continuous and binary target traits simultaneously. We demonstrate that PRSice-2 is dramatically faster and more memory-efficient than PRSice-1 and alternative PRS software, LDpred and lassosum, while having comparable predictive power.ConclusionPRSice-2's combination of efficiency and power will be increasingly important as data sizes grow and as the applications of PRS become more sophisticated, e.g., when incorporated into high-dimensional or gene set–based analyses. PRSice-2 is written in C++, with an R script for plotting, and is freely available for download from http://PRSice.info.

Similar Papers
  • Research Article
  • Cite Count Icon 471
  • 10.1086/522036
So Many Correlated Tests, So Little Time! Rapid Adjustment of P Values for Multiple Correlated Tests
  • Dec 1, 2007
  • American journal of human genetics
  • Karen N Conneely + 1 more

So Many Correlated Tests, So Little Time! Rapid Adjustment of P Values for Multiple Correlated Tests

  • Research Article
  • Cite Count Icon 19
  • 10.1093/gigascience/giab047
BIGwas: Single-command quality control and association testing for multi-cohort and biobank-scale GWAS/PheWAS data
  • Jun 29, 2021
  • GigaScience
  • Jan Christian Kässens + 2 more

BackgroundGenome-wide association studies (GWAS) and phenome-wide association studies (PheWAS) involving 1 million GWAS samples from dozens of population-based biobanks present a considerable computational challenge and are carried out by large scientific groups under great expenditure of time and personnel. Automating these processes requires highly efficient and scalable methods and software, but so far there is no workflow solution to easily process 1 million GWAS samples.ResultsHere we present BIGwas, a portable, fully automated quality control and association testing pipeline for large-scale binary and quantitative trait GWAS data provided by biobank resources. By using Nextflow workflow and Singularity software container technology, BIGwas performs resource-efficient and reproducible analyses on a local computer or any high-performance compute (HPC) system with just 1 command, with no need to manually install a software execution environment or various software packages. For a single-command GWAS analysis with 974,818 individuals and 92 million genetic markers, BIGwas takes ∼16 days on a small HPC system with only 7 compute nodes to perform a complete GWAS QC and association analysis protocol. Our dynamic parallelization approach enables shorter runtimes for large HPCs.ConclusionsResearchers without extensive bioinformatics knowledge and with few computer resources can use BIGwas to perform multi-cohort GWAS with 1 million GWAS samples and, if desired, use it to build their own (genome-wide) PheWAS resource. BIGwas is freely available for download from http://github.com/ikmb/gwas-qc and http://github.com/ikmb/gwas-assoc.

  • Research Article
  • Cite Count Icon 8
  • 10.1017/s1751731107000912
Bayesian prediction of breeding values for multivariate binary and continuous traits in simulated horse populations using threshold–linear models with Gibbs sampling
  • Jan 1, 2008
  • Animal
  • K.F Stock + 2 more

Bayesian prediction of breeding values for multivariate binary and continuous traits in simulated horse populations using threshold–linear models with Gibbs sampling

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 454
  • 10.1074/jbc.m703759200
Genome-scale Reconstruction of Metabolic Network in Bacillus subtilis Based on High-throughput Phenotyping and Gene Essentiality Data
  • Sep 1, 2007
  • Journal of Biological Chemistry
  • You-Kwan Oh + 4 more

In this report, a genome-scale reconstruction of Bacillus subtilis metabolism and its iterative development based on the combination of genomic, biochemical, and physiological information and high-throughput phenotyping experiments is presented. The initial reconstruction was converted into an in silico model and expanded in a four-step iterative fashion. First, network gap analysis was used to identify 48 missing reactions that are needed for growth but were not found in the genome annotation. Second, the computed growth rates under aerobic conditions were compared with high-throughput phenotypic screen data, and the initial in silico model could predict the outcomes qualitatively in 140 of 271 cases considered. Detailed analysis of the incorrect predictions resulted in the addition of 75 reactions to the initial reconstruction, and 200 of 271 cases were correctly computed. Third, in silico computations of the growth phenotypes of knock-out strains were found to be consistent with experimental observations in 720 of 766 cases evaluated. Fourth, the integrated analysis of the large-scale substrate utilization and gene essentiality data with the genome-scale metabolic model revealed the requirement of 80 specific enzymes (transport, 53; intracellular reactions, 27) that were not in the genome annotation. Subsequent sequence analysis resulted in the identification of genes that could be putatively assigned to 13 intracellular enzymes. The final reconstruction accounted for 844 open reading frames and consisted of 1020 metabolic reactions and 988 metabolites. Hence, the in silico model can be used to obtain experimentally verifiable hypothesis on the metabolic functions of various genes.

  • Research Article
  • Cite Count Icon 5
  • 10.1016/j.ajhg.2023.03.010
Scalable mixed model methods for set-based association studies on large-scale categorical data analysis and its application to exome-sequencing data in UK Biobank
  • Apr 4, 2023
  • The American Journal of Human Genetics
  • Wenjian Bi + 5 more

Scalable mixed model methods for set-based association studies on large-scale categorical data analysis and its application to exome-sequencing data in UK Biobank

  • PDF Download Icon
  • Supplementary Content
  • Cite Count Icon 7
  • 10.3389/fpls.2022.935748
The field phenotyping platform's next darling: Dicotyledons
  • Aug 24, 2022
  • Frontiers in Plant Science
  • Xiuni Li + 8 more

The genetic information and functional properties of plants have been further identified with the completion of the whole-genome sequencing of numerous crop species and the rapid development of high-throughput phenotyping technologies, laying a suitable foundation for advanced precision agriculture and enhanced genetic gains. Collecting phenotypic data from dicotyledonous crops in the field has been identified as a key factor in the collection of large-scale phenotypic data of crops. On the one hand, dicotyledonous plants account for 4/5 of all angiosperm species and play a critical role in agriculture. However, their morphology is complex, and an abundance of dicot phenotypic information is available, which is critical for the analysis of high-throughput phenotypic data in the field. As a result, the focus of this paper is on the major advancements in ground-based, air-based, and space-based field phenotyping platforms over the last few decades and the research progress in the high-throughput phenotyping of dicotyledonous field crop plants in terms of morphological indicators, physiological and biochemical indicators, biotic/abiotic stress indicators, and yield indicators. Finally, the future development of dicots in the field is explored from the perspectives of identifying new unified phenotypic criteria, developing a high-performance infrastructure platform, creating a phenotypic big data knowledge map, and merging the data with those of multiomic techniques.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 48
  • 10.3389/fgene.2019.00272
A QTL for Number of Teats Shows Breed Specific Effects on Number of Vertebrae in Pigs: Bridging the Gap Between Molecular and Quantitative Genetics
  • Mar 26, 2019
  • Frontiers in Genetics
  • Maren Van Son + 8 more

Modern breeding schemes for livestock species accumulate a large amount of genotype and phenotype data which can be used for genome-wide association studies (GWAS). Many chromosomal regions harboring effects on quantitative traits have been reported from these studies, but the underlying causative mutations remain mostly undetected. In this study, we combine large genotype and phenotype data available from a commercial pig breeding scheme for three different breeds (Duroc, Landrace, and Large White) to pinpoint functional variation for a region on porcine chromosome 7 affecting number of teats (NTE). Our results show that refining trait definition by counting number of vertebrae (NVE) and ribs (RIB) helps to reduce noise from other genetic variation and increases heritability from 0.28 up to 0.62 NVE and 0.78 RIB in Duroc. However, in Landrace, the effect of the same QTL on NTE mainly affects NVE and not RIB, which is reflected in reduced heritability for RIB (0.24) compared to NVE (0.59). Further, differences in allele frequencies and accuracy of rib counting influence genetic parameters. Correction for the top SNP does not detect any other QTL effect on NTE, NVE, or RIB in Landrace or Duroc. At the molecular level, haplotypes derived from 660K SNP data detects a core haplotype of seven SNPs in Duroc. Sequence analysis of 16 Duroc animals shows that two functional mutations of the Vertnin (VRTN) gene known to increase number of thoracic vertebrae (ribs) reside on this haplotype. In Landrace, the linkage disequilibrium (LD) extends over a region of more than 3 Mb also containing both VRTN mutations. Here, other modifying loci are expected to cause the breed-specific effect. Additional variants found on the wildtype haplotype surrounding the VRTN region in all sequenced Landrace animals point toward breed specific differences which are expected to be present also across the whole genome. This Landrace specific haplotype contains two missense mutations in the ABCD4 gene, one of which is expected to have a negative effect on the protein function. Together, the integration of largescale genotype, phenotype and sequence data shows exemplarily how population parameters are influenced by underlying variation at the molecular level.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 7
  • 10.1093/bib/bbaa033
GenoPheno: cataloging large-scale phenotypic and next-generation sequencing data within human datasets
  • Apr 6, 2020
  • Briefings in Bioinformatics
  • Alba Gutiérrez-Sacristán + 5 more

Precision medicine promises to revolutionize treatment, shifting therapeutic approaches from the classical one-size-fits-all to those more tailored to the patient’s individual genomic profile, lifestyle and environmental exposures. Yet, to advance precision medicine’s main objective—ensuring the optimum diagnosis, treatment and prognosis for each individual—investigators need access to large-scale clinical and genomic data repositories. Despite the vast proliferation of these datasets, locating and obtaining access to many remains a challenge. We sought to provide an overview of available patient-level datasets that contain both genotypic data, obtained by next-generation sequencing, and phenotypic data—and to create a dynamic, online catalog for consultation, contribution and revision by the research community. Datasets included in this review conform to six specific inclusion parameters that are: (i) contain data from more than 500 human subjects; (ii) contain both genotypic and phenotypic data from the same subjects; (iii) include whole genome sequencing or whole exome sequencing data; (iv) include at least 100 recorded phenotypic variables per subject; (v) accessible through a website or collaboration with investigators and (vi) make access information available in English. Using these criteria, we identified 30 datasets, reviewed them and provided results in the release version of a catalog, which is publicly available through a dynamic Web application and on GitHub. Users can review as well as contribute new datasets for inclusion (Web: https://avillachlab.shinyapps.io/genophenocatalog/; GitHub: https://github.com/hms-dbmi/GenoPheno-CatalogShiny).

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 6
  • 10.1093/database/baab051
A PostgreSQL Tripal solution for large-scale genotypic and phenotypic data
  • Aug 14, 2021
  • Database: The Journal of Biological Databases and Curation
  • Lacey-Anne Sanderson + 3 more

Researchers are seeking cost-effective solutions for management and analysis of large-scale genotypic and phenotypic data. Open-source software is uniquely positioned to fill this need through user-focused, crowd-sourced development. Tripal, an open-source toolkit for developing biological data web portals, uses the GMOD Chado database schema to achieve flexible, ontology-driven storage in PostgreSQL. Tripal also aids research-focused web portals in providing data according to findable, accessible, interoperable, reusable (FAIR) principles. We describe here a fully relational PostgreSQL solution to handle large-scale genotypic and phenotypic data that is implemented as a collection of freely available, open-source modules. These Tripal extension modules provide a holistic approach for importing, storage, display and analysis within a relational database schema. Furthermore, they embody the Tripal approach to FAIR data by providing multiple search tools and ensuring metadata is fully described and interoperable. Our solution focuses on data integrity, as well as optimizing performance to provide a fully functional system that is currently being used in the production of Tripal portals for crop species. We fully describe the implementation of our solution and discuss why a PostgreSQL-powered web portal provides an efficient environment for researcher-driven genotypic and phenotypic data analysis.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 646
  • 10.1038/s41588-021-00954-4
A generalized linear mixed model association tool for biobank-scale data.
  • Nov 1, 2021
  • Nature Genetics
  • Longda Jiang + 3 more

Compared with linear mixed model-based genome-wide association (GWA) methods, generalized linear mixed model (GLMM)-based methods have better statistical properties when applied to binary traits but are computationally much slower. In the present study, leveraging efficient sparse matrix-based algorithms, we developed a GLMM-based GWA tool, fastGWA-GLMM, that is severalfold to orders of magnitude faster than the state-of-the-art tools when applied to the UK Biobank (UKB) data and scalable to cohorts with millions of individuals. We show by simulation that the fastGWA-GLMM test statistics of both common and rare variants are well calibrated under the null, even for traits with extreme case-control ratios. We applied fastGWA-GLMM to the UKB data of 456,348 individuals, 11,842,647 variants and 2,989 binary traits (full summary statistics available at http://fastgwa.info/ukbimpbin ), and identified 259 rare variants associated with 75 traits, demonstrating the use of imputed genotype data in a large cohort to discover rare variants for binary complex traits.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 10
  • 10.1186/1471-2105-10-139
ECOMPAGT – efficient Combination and Management of Phenotypes and Genotypes for Genetic Epidemiology
  • May 11, 2009
  • BMC Bioinformatics
  • Sebastian Schönherr + 5 more

BackgroundHigh-throughput genotyping and phenotyping projects of large epidemiological study populations require sophisticated laboratory information management systems. Most epidemiological studies include subject-related personal information, which needs to be handled with care by following data privacy protection guidelines. In addition, genotyping core facilities handling cooperative projects require a straightforward solution to monitor the status and financial resources of the different projects.DescriptionWe developed a database system for an efficient combination and management of phenotypes and genotypes (eCOMPAGT) deriving from genetic epidemiological studies. eCOMPAGT securely stores and manages genotype and phenotype data and enables different user modes with different rights. Special attention was drawn on the import of data deriving from TaqMan and SNPlex genotyping assays. However, the database solution is adjustable to other genotyping systems by programming additional interfaces. Further important features are the scalability of the database and an export interface to statistical software.ConclusioneCOMPAGT can store, administer and connect phenotype data with all kinds of genotype data and is available as a downloadable version at .

  • Research Article
  • 10.12688/wellcomeopenres.23009.2
Implementation of a genotyped African population cohort, with virtual follow-up: A feasibility study in the Western Cape Province, South Africa.
  • Jan 13, 2025
  • Wellcome open research
  • Tsaone Tamuhla + 9 more

There is limited knowledge regarding African genetic drivers of disease due to prohibitive costs of large-scale genomic research in Africa. We piloted a scalable virtual genotyped cohort in South Africa that was affordable in this resource-limited context, cost-effective, scalable virtual genotyped cohort in South Africa, with participant recruitment using a tiered informed consent model and DNA collection by buccal swab. Genotype data was generated using the H3Africa Illumina micro-array, and phenotype data was derived from routine health data of participants. We demonstrated feasibility of nested case control genome wide association studies using these data for phenotypes type 2 diabetes mellitus (T2DM) and severe COVID-19. 2267346 variants were analysed in 459 participant samples, of which 229 (66.8%) are female. 78.6% of SNPs and 74% of samples passed quality control (QC). Principal component analysis showed extensive ancestry admixture in study participants. Of the 343 samples that passed QC, 93 participants had T2DM and 63 had severe COVID-19. For 1780 previously published COVID-19-associated variants, 3 SNPs in the pre-imputation data and 23 SNPS in the imputed data were significantly associated with severe COVID-19 cases compared to controls (p<0.05). For 2755 published T2DM associated variants, 69 SNPs in the pre-imputation data and 419 SNPs in the imputed data were significantly associated with T2DM cases when compared to controls (p<0.05). The results shown here are illustrative of what will be possible as the cohort expands in the future. Here we demonstrate the feasibility of this approach, recognising that the findings presented here are preliminary and require further validation once we have a sufficient sample size to improve statistical significance of findings.We implemented a genotyped population cohort with virtual follow up data in a resource-constrained African environment, demonstrating feasibility for scale up and novel health discoveries through nested case-control studies.

  • Research Article
  • 10.12688/wellcomeopenres.23009.1
Implementation of a genotyped African population cohort, with virtual follow-up: A feasibility study in the Western Cape Province, South Africa
  • Oct 22, 2024
  • Wellcome Open Research
  • Tsaone Tamuhla + 9 more

Background There is limited knowledge regarding African genetic drivers of disease due to prohibitive costs of large-scale genomic research in Africa. Methods We piloted a cost-effective, scalable virtual genotyped cohort in South Africa, with participant recruitment using a tiered informed consent model and DNA collection by buccal swab. Genotype data was generated using the H3Africa Illumina micro-array, and phenotype data was derived from routine health data of participants. We demonstrated feasibility of nested case control genome wide association studies using these data for phenotypes type 2 diabetes mellitus (T2DM) and severe COVID-19. Results 2267346 variants were analysed in 459 participant samples. 78.6% of SNPs and 74% of samples passed quality control (QC). Principal component analysis showed extensive ancestry admixture in study participants. For 1780 published COVID-19-associated variants, 3 SNPs in the pre-imputation data and 23 SNPS in the imputed data were significantly associated with severe COVID-19 cases compared to controls. For 2755 published T2DM associated variants, 69 SNPs in the pre-imputation data and 419 SNPs in the imputed data were significantly associated with T2DM cases when compared to controls. Conclusions The results shown here are illustrative of what will be possible as the cohort expands in the future. Here we demonstrate the feasibility of this approach, recognising that the findings presented here are preliminary and require further validation once we have a sufficient sample size to improve statistical significance of findings. We implemented a genotyped population cohort with virtual follow up data in a resource-constrained African environment, demonstrating feasibility for scale up and novel health discoveries through nested case-control studies.

  • Research Article
  • Cite Count Icon 12
  • 10.1177/14034948211004421
Integrating data from multiple Finnish biobanks and national health-care registers for retrospective studies: Practical experiences.
  • Apr 12, 2021
  • Scandinavian Journal of Public Health
  • Jaakko Lähteenmäki + 6 more

Aim: This case study aimed to investigate the process of integrating resources of multiple biobanks and health-care registers, especially addressing data permit application, time schedules, co-operation of stakeholders, data exchange and data quality. Methods: We investigated the process in the context of a retrospective study: Pharmacogenomics of antithrombotic drugs (PreMed study). The study involved linking the genotype data of three Finnish biobanks (Auria Biobank, Helsinki Biobank and THL Biobank) with register data on medicine dispensations, health-care encounters and laboratory results. Results: We managed to collect a cohort of 7005 genotyped individuals, thereby achieving the statistical power requirements of the study. The data collection process took 16 months, exceeding our original estimate by seven months. The main delays were caused by the congested data permit approval service to access national register data on health-care encounters. Comparison of hospital data lakes and national registers revealed differences, especially concerning medication data. Genetic variant frequencies were in line with earlier data reported for the European population. The yearly number of international normalised ratio (INR) tests showed stable behaviour over time. Conclusions: A large cohort, consisting of versatile individual-level phenotype and genotype data, can be constructed by integrating data from several biobanks and health data registers in Finland. Co-operation with biobanks is straightforward. However, long time periods need to be reserved when biobank resources are linked with national register data. There is a need for efforts to define general, harmonised co-operation practices and data exchange methods for enabling efficient collection of data from multiple sources.

  • Research Article
  • Cite Count Icon 14
  • 10.1016/j.asoc.2023.111128
Spatial-temporal traffic data imputation based on dynamic multi-level generative adversarial networks for urban governance
  • Dec 9, 2023
  • Applied Soft Computing
  • Bo Zhang + 2 more

Spatial-temporal traffic data imputation based on dynamic multi-level generative adversarial networks for urban governance

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant