Using probabilistic estimation of expression residuals (PEER) to obtain increased power and interpretability of gene expression analyses
We present PEER (probabilistic estimation of expression residuals), a software package implementing statistical models that improve the sensitivity and interpretability of genetic associations in population-scale expression data. This approach builds on factor analysis methods that infer broad variance components in the measurements. PEER takes as input transcript profiles and covariates from a set of individuals, and then outputs hidden factors that explain much of the expression variability. Optionally, these factors can be interpreted as pathway or transcription factor activations by providing prior information about which genes are involved in the pathway or targeted by the factor. The inferred factors are used in genetic association analyses. First, they are treated as additional covariates, and are included in the model to increase detection power for mapping expression traits. Second, they are analyzed as phenotypes themselves to understand the causes of global expression variability. PEER extends previous related surrogate variable models and can be implemented within hours on a desktop computer.
- Research Article
36
- 10.1186/1471-2105-9-557
- Dec 1, 2008
- BMC Bioinformatics
BackgroundResearchers wishing to conduct genetic association analysis involving single nucleotide polymorphisms (SNPs) or haplotypes are often confronted with the lack of user-friendly graphical analysis tools, requiring sophisticated statistical and informatics expertise to perform relatively straightforward tasks. Tools, such as the SimHap package for the R statistics language, provide the necessary statistical operations to conduct sophisticated genetic analysis, but lacks a graphical user interface that allows anyone but a professional statistician to effectively utilise the tool.ResultsWe have developed SimHap GUI, a cross-platform integrated graphical analysis tool for conducting epidemiological, single SNP and haplotype-based association analysis. SimHap GUI features a novel workflow interface that guides the user through each logical step of the analysis process, making it accessible to both novice and advanced users. This tool provides a seamless interface to the SimHap R package, while providing enhanced functionality such as sophisticated data checking, automated data conversion, and real-time estimations of haplotype simulation progress.ConclusionSimHap GUI provides a novel, easy-to-use, cross-platform solution for conducting a range of genetic and non-genetic association analyses. This provides a free alternative to commercial statistics packages that is specifically designed for genetic association analysis.
- Research Article
9
- 10.1101/pdb.top067819
- Feb 1, 2012
- Cold Spring Harbor Protocols
The approaches to identifying genes and genomic regions associated with human disease can be grouped into two categories: linkage analysis and genetic association analysis. Linkage analysis is useful for diseases of high penetrance that run strongly within families, but is limited in its ability to detect situations where there are multiple genes with smaller effects. An alternative is genetic association studies, which were initially performed on small numbers of candidate genes. This approach identified relatively few genes that were consistently associated with disease, but it is now possible to do a genetic association for the whole genome, making this approach more powerful. In practice, the two types of analysis are often interlinked. This article provides information on the tools needed to perform both genetic linkage and genetic association analysis.
- Research Article
13
- 10.1016/j.ajhg.2025.03.004
- Apr 1, 2025
- American journal of human genetics
Opportunities and challenges of local ancestry in genetic association analyses.
- Research Article
5
- 10.1186/s12859-025-06207-z
- Aug 19, 2025
- BMC bioinformatics
Genetic association studies play a pivotal role in identifying disease-associated variants, but researchers face challenges in performing essential calculations like Hardy-Weinberg equilibrium testing, odds ratios, and confidence intervals due to reliance on manual methods or multiple software tools. We aimed to develop GeneRiskCalc, an integrated web-based platform that simplifies genetic association analysis by automating Hardy-Weinberg equilibrium assessment, odds ratios with confidence interval calculation, and visual data presentation in case-control studies. Using an HTML/CSS/JavaScript framework, we developed online software with three core functionalities: (1) automated HWE evaluation, (2) odds ratio with 95% confidence interval computation with statistical validation, and (3) dynamic Forest Plot generation for data visualization. The tool was designed with an intuitive interface to minimize prerequisite statistical expertise. The tool, named the Genetic Risk Association Calculator (GeneRiskCalc), demonstrated high computational accuracy in HWE testing (χ2 validation) and association metrics (odds ratio and confidence interval). The results were cross-validated against established statistical methods, confirming their reliability. Furthermore, the integrated Forest Plotter enabled immediate visualization of effect sizes across multiple genetic models, facilitating a comprehensive interpretation of genetic associations. By integrating essential analytical steps into a single platform, the GeneRiskCalc, streamlines genetic epidemiology workflows, addressing key challenges in data analysis. Its user-friendly interface enhances accessibility, promotes reproducibility, and accelerates research in genetic association studies. The tool is freely available at GeneRiskCalc ( https://sites.google.com/view/GeneRiskCalc/home?authuser=0 ).
- Supplementary Content
- 10.17169/refubium-10852
- Jan 1, 2014
- Refubium (Universitätsbibliothek der Freien Universität Berlin)
Introduction: In the Berlin Aging Study II (BASE-II) a genomwide association study (GWAS) will be performed to examine the association of polymorphisms with special lipid parameters. The goal of the present study was to recruit a cohort which could be used to replicate GWAS data from BASE-II. In order to validate this study concentrated on three potential SNPs located outside of LPA on chromosome 6 which were genotyped as being under high suspicion to influence Lp(a) levels. To validate the used methods we tested two SNPs on LPA with a well known association. Additional to the genetic analysis a characterization of the study population especially for cardiovascular risk was performed. Methods: 500 patients of the outpatient department of lipid metabolism of the Charite-Universitatsmedizin Berlin were selected. All five SNPs were genotyped at Max-Planck-Institute for Molecular Genetics using commercially available assays based on TaqMan® chemistry following the manufacturer’s recommendations. Association analyses were carried out using the software PLINK v 1.07 and Lp(a) plasma levels as quantitative traits in an additive linear model, adjusted for age, sex, body mass index and apolipoprotein B plasma levels. Results: According to the guidelines of the German Cardiac Society the analysis of cardiovascular risk showed that 60 % of our study population is classified as high risk subjects. The genotype analysis demonstrated a highly significant association for SNP rs10455872 on LPA with Lp(a) levels. There were no associations for three additional SNPs on LPA, TNFRSF11A on chromosome 18 and TFPI on chromosome 2. rs17210569, located near TRPC4 on chromosome 13, showed no significant association ( p = 0,08274 ), either. Regarding previous results these analyses indicate an association of rs17210569 with Lp(a) levels. Conclusions: We successfully validated our replication cohort and created an acceptable basis for further genetic association analysis with lipid parameters of the BASE-II cohort. Our study corroborates previous evidence regarding the involvement of rs17210569 near TRPC4 on chromosome 13 in controlling Lp(a) plasma levels. The repetition of these results in a bigger study population is in process of planning.
- Research Article
12
- 10.1038/s41746-023-00903-x
- Aug 21, 2023
- npj Digital Medicine
Electronic health records are often incomplete, reducing the power of genetic association studies. For some diseases, such as knee osteoarthritis where the routine course of diagnosis involves an X-ray, image-based phenotyping offers an alternate and unbiased way to ascertain disease cases. We investigated this by training a deep-learning model to ascertain knee osteoarthritis cases from knee DXA scans that achieved clinician-level performance. Using our model, we identified 1931 (178%) more cases than currently diagnosed in the health record. Individuals diagnosed as cases by our model had higher rates of self-reported knee pain, for longer durations and with increased severity compared to control individuals. We trained another deep-learning model to measure the knee joint space width, a quantitative phenotype linked to knee osteoarthritis severity. In performing genetic association analysis, we found that use of a quantitative measure improved the number of genome-wide significant loci we discovered by an order of magnitude compared with our binary model of cases and controls despite the two phenotypes being highly genetically correlated. In addition we discovered associations between our quantitative measure of knee osteoarthritis and increased risk of adult fractures- a leading cause of injury-related death in older individuals-, illustrating the capability of image-based phenotyping to reveal epidemiological associations not captured in the electronic health record. For diseases with radiographic diagnosis, our results demonstrate the potential for using deep learning to phenotype at biobank scale, improving power for both genetic and epidemiological association analysis.
- Research Article
17
- 10.1002/gepi.20197
- Dec 22, 2006
- Genetic Epidemiology
Genotype misclassification occurs frequently in human genetic association studies. When cases and controls are subject to the same misclassification model, Pearson's chi-square test has the correct type I error but may lose power. Most current methods adjusting for genotyping errors assume that the misclassification model is known a priori or can be assessed by a gold standard instrument. But in practical applications, the misclassification probabilities may not be completely known or the gold standard method can be too costly to be available. The repeated measurement design provides an alternative approach for identifying misclassification probabilities. With this design, a proportion of the subjects are measured repeatedly (five or more repeats) for the genotypes when the error model is completely unknown. We investigate the applications of the repeated measurement method in genetic association analysis. Cost-effectiveness study shows that if the phenotyping-to-genotyping cost ratio or the misclassification rates are relatively large, the repeat sampling can gain power over the regular case-control design. We also show that the power gain is not sensitive to the genetic model, genetic relative risk and the population high-risk allele frequency, all of which are typically important ingredients in association studies. An important implication of this result is that whatever the genetic factors are, the repeated measurement method can be applied if the genotyping errors must be accounted for or the phenotyping cost is high.
- Conference Article
1
- 10.1109/bibmw.2012.6470249
- Oct 1, 2012
Gene-gene interactions are important factors underlying a common complex trait that is mostly polygenic. While many methods have been proposed to analyze gene-gene interactions in genetic association studies, the interpretation of the identified gene-gene interactions is not straightforward. In order to aid the interpretation of gene-gene interactions, we developed the GxG-Viztool, an executable program for visualizing gene-gene interactions in genetic association analysis. The GxG-Viztool provides an effective way to recognize genotype combinations that enhance/repress a trait and to display polygenic structure of interactions. The GxG-Viztool implements six graphical tools: checkerboard, pairwise checkerboard, forest, funnel, 3D lattice and parallel coordinate plots, which make it effective to recognize certain patterns in gene-gene interactions. It is freely available at http ://bibs. snu. ac. kr/GxG-Viz tool.
- Research Article
101
- 10.1186/1471-2105-9-290
- Jun 23, 2008
- BMC Bioinformatics
BackgroundSince the completion of the HapMap project, huge numbers of individual genotypes have been generated from many kinds of laboratories. The efforts of finding or interpreting genetic association between disease and SNPs/haplotypes have been on-going widely. So, the necessity of the capability to analyze huge data and diverse interpretation of the results are growing rapidly.ResultsWe have developed an advanced tool to perform linkage disequilibrium analysis, and genetic association analysis between disease and SNPs/haplotypes in an integrated web interface. It comprises of four main analysis modules: (i) data import and preprocessing, (ii) haplotype estimation, (iii) LD blocking and (iv) association analysis. Hardy-Weinberg Equilibrium test is implemented for each SNPs in the data preprocessing. Haplotypes are reconstructed from unphased diploid genotype data, and linkage disequilibrium between pairwise SNPs is computed and represented by D', r2 and LOD score. Tagging SNPs are determined by using the square of Pearson's correlation coefficient (r2). If genotypes from two different sample groups are available, diverse genetic association analyses are implemented using additive, codominant, dominant and recessive models. Multiple verified algorithms and statistics are implemented in parallel for the reliability of the analysis.ConclusionSNPAnalyzer 2.0 performs linkage disequilibrium analysis and genetic association analysis in an integrated web interface using multiple verified algorithms and statistics. Diverse analysis methods, capability of handling huge data and visual comparison of analysis results are very comprehensive and easy-to-use.
- Abstract
- 10.1016/j.healun.2022.01.1740
- Apr 1, 2022
- The Journal of Heart and Lung Transplantation
Whole-Genome Sequencing Reveals Genetic Variant at the HLA Locus Associated with Graft Failure After Heart Transplantation
- Research Article
- 10.1002/brb3.70848
- Sep 1, 2025
- Brain and Behavior
ABSTRACTBackgroundGrowing evidence suggests a close association between circulating micronutrient levels and neuroimmune diseases. Nevertheless, the causal relationship between them remains unclear. Furthermore, due to confounding factors, many micronutrients implicated in these diseases remain unidentified. This study aimed to determine the causal relationship between circulating micronutrients and neuroimmune diseases through genetic association analysis, and to analyze the regulatory role of circulating micronutrients in neuroimmune diseases.MethodIn this study, we used a two‐sample mendelian randomization (MR) analysis to explore the causal relationship between micronutrients levels and neuroimmune disease. Fourteen micronutrients were screened from a published genome‐wide association study (GWAS). Neuroimmune diseases include multiple sclerosis (MS), Guillain–Barre syndrome (GBS), acute disseminated encephalomyelitis (ADEM), acute poliomyelitis (AP), sequelae of poliomyelitis (SP), optic neuritis (ON), and myasthenia gravis (MG). Data on these seven neuroimmune diseases came from the FinnGen database and included 5523 cases and 2,860,006 controls. The inverse variance weighting (IVW) method was used as the main MR analysis method, and sensitivity analysis was performed to determine MR hypotheses.ResultsThrough MR analysis and sensitivity testing, we identified significant causal relationships between four neuroimmune diseases and micronutrient levels. Specifically, MS was causally associated with magnesium levels (OR: 0.467, 95% CI: 0.269–0.809, p = 0.007), ADEM with folate levels (OR: 0.022, 95% CI: 0.001–0.957, p = 0.047), ON with vitamin B6 levels (OR: 0.382, 95% CI: 0.187–0.778, p = 0.008), and MG with iron levels (OR: 0.194, 95% CI: 0.043–0.867, p = 0.032). Sensitivity analysis showed that there was no level pleiotropic or heterogeneity in our study results.ConclusionThis study established the causal relationship between micronutrients and neuroimmune diseases. These findings provide new insights into the etiology of neuroimmune diseases and provide a theoretical basis for micronutrient regulation, prevention, and treatment of neuroimmune diseases.
- Research Article
28
- 10.1093/aje/kwp180
- Jul 27, 2009
- American Journal of Epidemiology
In genetic association studies, investigators compare allele or genotype frequencies in unrelated case and control subjects or examine preferential allele transmissions from parents to affected offspring. In many genetic case-control studies, the collection of DNA material extends to relatives such as parents of cases. Thus, case-control and case-parent trio association analyses are possible. Whereas the goal of collecting genetic information from family members in a study initially designed as a case-control study is to enrich the genetic analysis, increase power, or address concern about population structure bias, methods of combining genetic data from unrelated case and control subjects with genetic trio data from the same study population are not well known. A number of hybrid approaches have been developed that utilize such data together. In this paper, the authors describe key features of genetic case-control and case-parent trio studies and review commonly used methods of genetic analysis for case-parent trio designs. In addition, they provide a pragmatic review of statistical methods and available software for existing hybrid approaches that combine various components of case-control and genetic trio data. The application of all methods is illustrated using a candidate gene study of childhood leukemia that included case-control subjects and their parents.
- Research Article
54
- 10.1002/gepi.20393
- Feb 13, 2009
- Genetic Epidemiology
In genetic association studies, different complex phenotypes are often associated with the same marker. Such associations can be indicative of pleiotropy (i.e. common genetic causes), of indirect genetic effects via one of these phenotypes, or can be solely attributable to non-genetic/environmental links between the traits. To identify the phenotypes with the inducing genetic association, statistical methodology is needed that is able to distinguish between the different causes of the genetic associations. Here, we propose a simple, general adjustment principle that can be incorporated into many standard genetic association tests which are then able to infer whether an SNP has a direct biological influence on a given trait other than through the SNP's influence on another correlated phenotype. Using simulation studies, we show that, in the presence of a non-marker related link between phenotypes, standard association tests without the proposed adjustment can be biased. In contrast to that, the proposed methodology remains unbiased. Its achieved power levels are identical to those of standard adjustment methods, making the adjustment principle universally applicable in genetic association studies. The principle is illustrated by an application to three genome-wide association analyses.
- Research Article
- 10.14288/1.0376773
- Jan 1, 2017
- Open Collections
Genetic association analysis across tree-structure routine healthcare data
- Research Article
- 10.4070/kcj.2006.36.10.688
- Jan 1, 2006
- Korean Circulation Journal
Background and Objectives:The common methods of genetic association analysis are sensitive to population stratification, which may easily lead to a spurious association result. We used a regression approach based for lin- kage disequilibrium to perform a high resolution genetic association analysis. Subjects and Methods:We applied a regression approach that can increase the resolution of quantitative traits that are related with cardiovascular di- seases. The population data was composed of 543 males and 876 females without cardiovascular diseases, and it was obtained from a cardiovascular genome center. We used information about linkage disequilibrium between the mar- ker and trait locus, and we added the covariates to model their effects. Results:We found that this regression approach has the merit of analyzing genetic association based on linkage disequilibrium. In the analysis of the male group, the total cholesterol was significantly in linkage disequilibrium with CETP3 (p=0.002), and triglyceride was significantly in linkage disequilibrium with ACE8 (p=0.037), APOA1-1 (p=0.031), APOA5-1 (p=0.001), APOA5- 2 (p=0.001) and LIPC4 (p=0.022). HDL-cholesterol was significantly in linkage disequilibrium with ACE7 (p= 0.002), ACE8 (p=0.008), ACE10 (p=0.003), APOA5-2 (p=0.022), and MTP1 (p=0.001). In the female group, total cholesterol was significantly associated with APOA5-1 (p=0.020), APOA5-2 (p=0.001), and LIPC1 (p= 0.016), and triglyceride was significantly associated with APOA5-1 (p=0.009), APOA5-2 (p=0.001), and CETP5 (p=0.049). LDL-cholesterol was significantly associated with APOA5-2 (p=0.004), and HDL-cholesterol was significantly associated with LIPC1 (p=0.004). Conclusion:We used a regression-based method to perform high resolution linkage disequilibrium analysis of a quantitative trait locus that's associated with lipid profiles. This me- thod of using a single marker, as applied in this paper, was well suited for analysis of genetic association. Because of the simplicity, the method can also be easily performed by routine statistical analysis software. (Korean Circulation J 2006;36:688-694)