Higher criticism to compare two large frequency tables, with sensitivity to possible rare and weak differences

David L Donoho,Alon Kipnis

doi:10.1214/21-aos2158

Abstract

We adapt Higher Criticism (HC) to the comparison of two frequency tables which may—or may not—exhibit moderate differences between the tables in some unknown, relatively small subset out of a large number of categories. Our analysis of the power of the proposed HC test quantifies the rarity and size of assumed differences and applies moderate deviations-analysis to determine the asymptotic powerfulness/powerlessness of our proposed HC procedure. Our analysis considers the null hypothesis of no difference in underlying generative model against a rare/weak perturbation alternative, in which the frequencies of N1−β out of the N categories are perturbed by r(logN)/2n in the Hellinger distance; here, n is the size of each sample. Our proposed Higher Criticism (HC) test for this setting uses P-values obtained from N exact binomial tests. We characterize the asymptotic performance of the HC-based test in terms of the sparsity parameter β and the perturbation intensity parameter r. Specifically, we derive a region in the (β,r)-plane where the test asymptotically has maximal power, while having asymptotically no power outside this region. Our analysis distinguishes between cases in which the counts in both tables are low, versus cases in which counts are high, corresponding to the cases of sparse and dense frequency tables. The phase transition curve of HC in the high-counts regime matches formally the curve delivered by HC in a two-sample normal means model.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Higher criticism to compare two large frequency tables, with sensitivity to possible rare and weak differences

Abstract

Talk to us

Similar Papers

More From: The Annals of Statistics

Lead the way for us

Journal: The Annals of Statistics	Publication Date: Jun 1, 2022
Citations: 8

Similar Papers

EBT: a statistic test identifying moderate size of significant features with balanced power and precision for genome-wide rate comparisons.
Xinjie Hui ... Rongfei Han
Bioinformatics | VOL. 33
Xinjie Hui, et. al.Xinjie Hui ... Rongfei Han
03 May 2017
Bioinformatics | VOL. 33

Comparison of Correspondence Analysis based on Hellinger and chi-square distances to obtain sensory spaces from check-all-that-apply (CATA) questions
Leticia Vidal ... Sara R Jaeger
Food Quality and Preference | VOL. 43
Leticia Vidal, et. al.Leticia Vidal ... Sara R Jaeger
11 Mar 2015
Food Quality and Preference | VOL. 43

Phase III efficacy and safety trial of a new leuprolide acetate 3.75 mg depot formulation in prostate cancer patients
...
Journal of Clinical Oncology | VOL. 27
, et. al. ...
20 May 2009
Journal of Clinical Oncology | VOL. 27

Cosmological Non-Gaussian Signature Detection: Comparing Performance of Different Statistical Tests
J Jin ... D L Donoho
EURASIP Journal on Advances in Signal Processing | VOL. 2005
J Jin, et. al.J Jin ... D L Donoho
14 Sep 2005
EURASIP Journal on Advances in Signal Processing | VOL. 2005

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Higher criticism to compare two large frequency tables, with sensitivity to possible rare and weak differences

Abstract

Talk to us

Similar Papers

More From: The Annals of Statistics