Genetic prediction in the Genetic Analysis Workshop 18 sequencing data.

Andreas Ziegler,Nora Bohossian,Vincent P Diego,Chen Yao

doi:10.1002/gepi.21826

Abstract

High-throughput sequencing data can be used to predict phenotypes from genotypes, and this corresponds to establishing a prognostic model. In extended pedigrees the relatedness of subjects provides additional information so that genetic values, fixed or random genetic components, and heritability can be estimated. At the Genetic Analysis Workshop 18, the working group on genetic prediction dealt with both establishing a prognostic model and, in one contribution, comparing standard logistic regression with robust logistic regression in a sample of unrelated affected or unaffected individuals. Results of both logistic regression approaches were similar. All other contributions to this group used extended family data, in general using the quantitative trait blood pressure. The individual contributions varied in several important aspects, such as the estimation of the kinship matrix and the estimation method. Contributors chose various approaches for model validation, including different versions of cross-validation or within-family validation. Within-family validation included model building in the upper generations and validation in later generations. The choice of the statistical model and the computational algorithm had substantial effects on computation time. If decorrelation approaches were applied, the computational burden was substantially reduced. Some software packages estimated negative eigenvalues, although eigenvalues of correlation matrices should be non-negative. Most statistical models and software packages have been developed for experimental crosses and planned breeding programs. With their specialized pedigree structures, they are not sufficiently flexible to accommodate the variability of human pedigrees in general, and improved implementations are required.

Full Text