Geometric consistency of principal component scores for high‐dimensional mixture models and its application

Kazuyoshi Yata,Makoto Aoshima

doi:10.1111/sjos.12432

Kazuyoshi Yata, Makoto Aoshima

Open Access

https://doi.org/10.1111/sjos.12432

Copy DOI

Journal: Scandinavian Journal of Statistics	Publication Date: Dec 23, 2019
Citations: 6	License type: CC BY 4.0

Affiliation: University of Tsukuba

Abstract

AbstractIn this article, we consider clustering based on principal component analysis (PCA) for high‐dimensional mixture models. We present theoretical reasons why PCA is effective for clustering high‐dimensional data. First, we derive a geometric representation of high‐dimension, low‐sample‐size (HDLSS) data taken from a two‐class mixture model. With the help of the geometric representation, we give geometric consistency properties of sample principal component scores in the HDLSS context. We develop ideas of the geometric representation and provide geometric consistency properties for multiclass mixture models. We show that PCA can cluster HDLSS data under certain conditions in a surprisingly explicit way. Finally, we demonstrate the performance of the clustering using gene expression datasets.

Full Text