Loss Of Statistical Power Research Articles

BackgroundCurrent methods of high-dimensional unsupervised clustering of mass cytometry data lack means to monitor and evaluate clustering results. Whether unsupervised clustering is correct is typically evaluated by agreement with dimensionality reduction techniques or based on benchmarking with manually classified cells. The ambiguity and lack of reproducibility of sequential gating has been replaced with ambiguity in interpretation of clustering results. On the other hand, spurious overclustering of data leads to loss of statistical power. We have developed INFLECT, an R-package designed to give insight in clustering results and provide an optimal number of clusters. In our approach, a mass cytometry dataset is overclustered intentionally to ensure the smallest phenotypically different subsets are captured using FlowSOM. A range of metacluster number endpoints are generated and evaluated using marker interquartile range and distribution unimodality checks. The fraction of marker distributions that pass these checks is taken as a measure of clustering success. The fraction of unimodal distributions within metaclusters is plotted against the number of generated metaclusters and reaches a plateau of diminishing returns. The inflection point at which this occurs gives an optimal point of capturing cellular heterogeneity versus statistical power.ResultsWe applied INFLECT to four publically available mass cytometry datasets of different size and number of markers. The unimodality score consistently reached a plateau, with an inflection point dependent on dataset size and number of dimensions. We tested both ConsenusClusterPlus metaclustering and hierarchical clustering. While hierarchical clustering is less computationally expensive and thus faster, it achieved similar results to ConsensusClusterPlus. The four datasets consisted of labeled data and we compared INFLECT metaclustering to published results. INFLECT identified a higher optimal number of metaclusters for all datasets. We illustrated the underlying heterogeneity within labels, showing that these labels encompass distinct types of cells.ConclusionINFLECT addresses a knowledge gap in high-dimensional cytometry analysis, namely assessing clustering results. This is done through monitoring marker distributions for interquartile range and unimodality across a range of metacluster numbers. The inflection point is the optimal trade-off between cellular heterogeneity and statistical power, applied in this work for FlowSOM clustering on mass cytometry datasets.

Read full abstract

The measurement of many human traits, states, and disorders begins with a set of items on a questionnaire. The response format for these questions is often simply binary (e.g., yes/no) or ordered (e.g., high, medium or low). During data analysis, these items are frequently summed or used to estimate factor scores. In clinical applications, such assessments are often non-normally distributed in the general population because many respondents are unaffected, and therefore asymptomatic. As a result, in many cases these measures violate the statistical assumptions required for subsequent analyses. To reduce the influence of the non-normality and quasi-continuous assessment, variables are frequently recoded into binary (affected-unaffected) or ordinal (mild-moderate-severe) diagnoses. Ordinal data therefore present challenges at multiple levels of analysis. Categorizing continuous variables into ordered categories typically results in a loss of statistical power, which represents an incentive to the data analyst to assume that the data are normally distributed, even when they are not. Despite prior zeitgeists suggesting that, e.g., variables with more than 10 ordered categories may be regarded as continuous and analyzed as if they were, we show via simulation studies that this is not generally the case. In particular, using Pearson product-moment correlations instead of maximum likelihood estimates of polychoric correlations biases the estimated correlations towards zero. This bias is especially severe when a plurality of the observations fall into a single observed category, such as a score of zero. By contrast, estimating the ordinal correlation by maximum likelihood yields no estimation bias, although standard errors are (appropriately) larger. We also illustrate how odds ratios depend critically on the proportion or prevalence of affected individuals in the population, and therefore are sub-optimal for studies where comparisons of association metrics are needed. Finally, we extend these analyses to the classical twin model and demonstrate that treating binary data as continuous will underestimate genetic and common environmental variance components, and overestimate unique environment (residual) variance. These biases increase as prevalence declines. While modeling ordinal data appropriately may be more computationally intensive and time consuming, failing to do so will likely yield biased correlations and biased parameter estimates from modeling them.

Read full abstract

Loss Of Statistical Power Research Articles

Articles published on Loss Of Statistical Power

INFLECT: an R-package for cytometry cluster evaluation using marker modality

Evaluation of machine learning methods for covariate data imputation in pharmacometrics.

The case against censoring of progression-free survival in cancer clinical trials – A pandemic shutdown as an illustration

Exploratory analyses of clinical trial data used for health technology assessments: a retrospective evaluation.

Reclaiming independence in spatial-clustering datasets: A series of data-driven spatial weights matrices.

Efficacy of a Toothpaste Based on Microcrystalline Hydroxyapatite on Children with Hypersensitivity Caused by MIH: A Randomised Controlled Trial.

Efficient spline regression for neural spiking data.

Widespread and interrelated gray matter reductions in child sexual offenders with and without pedophilia: Evidence from a multivariate structural MRI study.

Impact of Tumor Assessment Frequency on Statistical Power in Randomized Cancer Clinical Trials Evaluating Progression-Free Survival.

2dFDR: a new approach to confounder adjustment substantially increases detection power in omics association studies

The impact of left truncation of exposure in environmental case-control studies: evidence from breast cancer risk associated with airborne dioxin.

Analytic results for scalar-mediated Higgs boson production in association with two jets

Genotype imputation in case-only studies of gene-environment interaction: validity and power

Contour prognostic model for predicting survival after resection of colorectal liver metastases: development and multicentre validation study using largest diameter and number of metastases with RAS mutation status.

Impact of covariate omission and categorization from the Fine–Gray model in randomized-controlled trials

Best Practices for Binary and Ordinal Data Analyses.

Study on the Missing Data Mechanisms and Imputation Methods

IMIX: a multivariate mixture model approach to association analysis through multi-omics data integration.

DCTRGAN: improving the precision of generative models with reweighting

Hybrid of Restricted and Penalized Maximum Likelihood Method for Efficient Genome-Wide Association Study.

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Loss Of Statistical Power Research Articles

Articles published on Loss Of Statistical Power

INFLECT: an R-package for cytometry cluster evaluation using marker modality

Evaluation of machine learning methods for covariate data imputation in pharmacometrics.

The case against censoring of progression-free survival in cancer clinical trials – A pandemic shutdown as an illustration

Exploratory analyses of clinical trial data used for health technology assessments: a retrospective evaluation.

Reclaiming independence in spatial-clustering datasets: A series of data-driven spatial weights matrices.

Efficacy of a Toothpaste Based on Microcrystalline Hydroxyapatite on Children with Hypersensitivity Caused by MIH: A Randomised Controlled Trial.

Efficient spline regression for neural spiking data.

Widespread and interrelated gray matter reductions in child sexual offenders with and without pedophilia: Evidence from a multivariate structural MRI study.

Impact of Tumor Assessment Frequency on Statistical Power in Randomized Cancer Clinical Trials Evaluating Progression-Free Survival.

2dFDR: a new approach to confounder adjustment substantially increases detection power in omics association studies

The impact of left truncation of exposure in environmental case-control studies: evidence from breast cancer risk associated with airborne dioxin.

Analytic results for scalar-mediated Higgs boson production in association with two jets

Genotype imputation in case-only studies of gene-environment interaction: validity and power

Contour prognostic model for predicting survival after resection of colorectal liver metastases: development and multicentre validation study using largest diameter and number of metastases with RAS mutation status.

Impact of covariate omission and categorization from the Fine–Gray model in randomized-controlled trials

Best Practices for Binary and Ordinal Data Analyses.

Study on the Missing Data Mechanisms and Imputation Methods

IMIX: a multivariate mixture model approach to association analysis through multi-omics data integration.

DCTRGAN: improving the precision of generative models with reweighting

Hybrid of Restricted and Penalized Maximum Likelihood Method for Efficient Genome-Wide Association Study.