Convex relaxation for the generalized maximum-entropy sampling problem
Abstract The generalized maximum-entropy sampling problem (GMESP) is to select an order- s principal submatrix from an order- n covariance matrix, to maximize the product of its t greatest eigenvalues, $$0<t\le s <n$$ 0 < t ≤ s < n . Introduced more than 25 years ago, GMESP is a natural generalization of two fundamental problems in statistical design theory: (i) maximum-entropy sampling problem (MESP); (ii) binary D-optimality (D-Opt). In the general case, it can be motivated by a selection problem in the context of principal component analysis (PCA). We introduce the first convex-optimization based relaxation for GMESP, study its behavior, compare it to an earlier spectral bound, and demonstrate its use in a branch-and-bound scheme. We find that such an approach is practical when $$s-t$$ s - t is very small.
- Book Chapter
1663
- 10.1002/0471271357.ch12
- Feb 22, 2002
Principal component analysis is a one-sample technique applied to data with no groupings among the observations and no partitioning of the variables into subvectors y and x. Principal components are concerned only with the core structure of a single sample of observations on p variables. In principal component analysis, we seek to maximize the variance of a linear combination of the variables. For example, the first principal component could be used to rank students on the basis of their scores on achievement tests in English, mathematics, reading, and so on. An average score would provide a single scale on which to compare the students, but with unequal weights in the principal component, we can spread the students out further on the scale and obtain a better ranking. The first principal component is the linear combination with maximal variance. The second principal component is the linear combination with maximal variance in a direction orthogonal to the first principal component, and so on. The first principal component also represents the line that minimizes the total sum of squared perpendicular distances from the points to the line. Principal components are often used as a dimension reduction device to obtain a smaller number of variables for input to another analysis. Another useful dimension reduction device is to evaluate the first two principal components for each observation vector and construct a scatter plot to check for multivariate normality, outliers, and so on. The properties of principal components can be interpreted either geometrically or algebraically. Principal components are orthogonal because they are formed with eigenvectors of the covariance (or correlation) matrix, which is symmetric. Principal components are not scale invariant because a scale chance in one of the variables leads to a change in the shape of the swarm of points in the sample. Since principal components are not scale invariant, the components extracted from a covariance matrix differ from those obtained from the corresponding correlation matrix. In fact, the components of a given correlation matrix will serve for other correlation matrices. There are four common methods that can be used to decide how many components to retain in order to effectively summarize the data. Three of the four methods are based on the eigenvalues of the covariance matrix (or correlation matrix). Because of the lack of scale invariance of principal components from the covariance matrix, the coefficients cannot be converted to standardized form, as can be done with coefficients in discriminant functions in Chapter 8 and canonical variates in Chapter 11. Hence we use the coefficients themselves for interpretation. We must choose between the covariance matrix or the correlation matrix, knowing they will yield a different interpretation. One aid to interpretation is to note that for certain patterns of elements in the covariance matrix or the correlation matrix, the form of the principal components can be predicted. For example, if one variable has a much larger variance than the other variables, this variable will dominate the first component, which will account for most of the variance. Another case in which a component will duplicate a variable occurs when the variable is uncorrelated with the other variables. In some settings, the principal components can be interpreted as measures of size and shape. Various methods of selecting a subset of variables are discussed. Since there is no grouping variable or dependent variable in the setting of principal components, we wish to find the subst that best captures the internal variation and covariation of the variables. As usual, examples and problems amply illustrate the techniques in this chapter.
- Book Chapter
2
- 10.1007/978-1-4939-1507-1_7
- Jan 1, 2014
For multivariate analysis with p variables the problem that often arises is the ambiguous nature of the correlation or covariance matrix. When p is moderately or very large it is generally difficult to identify the true nature of relationship among the variables as well as observations from the covariance or correlation matrix. Under such situations a very common way to simplify the matter is to reduce the dimension by considering only those variables (actual or derived) which are truly responsible for the overall variation. Important and useful dimension reduction techniques are Principal Component Analysis (PCA), Factor Analysis, Multidimensional Scaling, Independent Component Analysis (ICA), etc. Among them PCA is the most popular one. One may look at this method in three different ways. It may be considered as a method of transforming correlated variables into uncorrelated one or a method of finding linear combinations with relatively small or large variability or a tool for data reduction. The third criterion is more data oriented. In PCA primarily it is not necessary to make any assumption regarding the underlying multivariate distribution but if we are interested in some inference problems related to PCA then assumption of multivariate normality is necessary. The eigen values and eigen vectors of the covariance or correlation matrix are the main contributors of a PCA. The eigen vectors determine the directions of maximum variability whereas the eigen values specify the variances. In practice, decisions regarding the quality of the principal component approximation should be made on the basis of eigen value–eigen vector pairs.
- Preprint Article
- 10.1103/030334
- Apr 7, 2022
Principal component analysis (PCA) is a dimensionality reduction method in data analysis that involves diagonalizing the covariance matrix of the dataset. Recently, quantum algorithms have been formulated for PCA based on diagonalizing a density matrix. These algorithms assume that the covariance matrix can be encoded in a density matrix, but a concrete protocol for this encoding has been lacking. Our work aims to address this gap. Assuming amplitude encoding of the data, with the data given by the ensemble $\{p_i,| \psi_i \rangle\}$, then one can easily prepare the ensemble average density matrix $\overline{\rho} = \sum_i p_i |\psi_i\rangle \langle \psi_i |$. We first show that $\overline{\rho}$ is precisely the covariance matrix whenever the dataset is centered. For quantum datasets, we exploit global phase symmetry to argue that there always exists a centered dataset consistent with $\overline{\rho}$, and hence $\overline{\rho}$ can always be interpreted as a covariance matrix. This provides a simple means for preparing the covariance matrix for arbitrary quantum datasets or centered classical datasets. For uncentered classical datasets, our method is so-called "PCA without centering", which we interpret as PCA on a symmetrized dataset. We argue that this closely corresponds to standard PCA, and we derive equations and inequalities that bound the deviation of the spectrum obtained with our method from that of standard PCA. We numerically illustrate our method for the MNIST handwritten digit dataset. We also argue that PCA on quantum datasets is natural and meaningful, and we numerically implement our method for molecular ground-state datasets.
- Research Article
147
- 10.1016/s0950-3293(01)00017-9
- Jun 20, 2001
- Food Quality and Preference
Principal component analysis in sensory analysis: covariance or correlation matrix?
- Book Chapter
3
- 10.5772/4844
- Jul 1, 2007
In this chapter, we have shown the class of low-rank approximation algorithms based directly on image data. In general, those algorithms are reduced to a couple of eigenvalue
- Research Article
- 10.1109/tit.2024.3514795
- Feb 1, 2025
- IEEE Transactions on Information Theory
Eigenvector perturbation analysis plays a vital role in various data science applications. A large body of prior works, however, focused on establishing <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$\ell _{2}$ </tex-math></inline-formula> eigenvector perturbation bounds, which are often highly inadequate in addressing tasks that rely on fine-grained behavior of an eigenvector. This paper makes progress on this by studying the perturbation of linear functions of an unknown eigenvector. Focusing on two fundamental problems — matrix denoising and principal component analysis — in the presence of Gaussian noise, we develop a suite of statistical theory that characterizes the perturbation of arbitrary linear functions of an unknown eigenvector. In order to mitigate a non-negligible bias issue inherent to the natural “plug-in” estimator, we develop de-biased estimators that <xref ref-type="disp-formula" rid="deqn1" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">(1)</xref> achieve minimax lower bounds for a family of scenarios (modulo some logarithmic factor), and <xref ref-type="disp-formula" rid="deqn2" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">(2)</xref> can be computed in a data-driven manner without sample splitting. Noteworthily, the proposed estimators are nearly minimax optimal even when the associated eigen-gap is <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">substantially smaller</i> than what is required in prior statistical theory.
- Research Article
37
- 10.1002/cem.2880
- Apr 1, 2017
- Journal of Chemometrics
The covariance matrix (or its inverse, the precision matrix) is central to many chemometric techniques. Traditional sample estimators perform poorly for high‐dimensional data such as metabolomics data. Because of this, many traditional inference techniques break down or produce unreliable results. In this paper, we selectively review several modern estimators of the covariance and precision matrix that improve upon the traditional sample estimator. We focus on 3 general techniques: eigenvalue‐shrinkage estimation, ridge‐type estimation, and structured estimation. These methods rely on different assumptions regarding the structure of the covariance or precision matrix. Various examples, in particular using metabolomics data, are used to compare these techniques and to demonstrate that in concert with, eg, principal component analysis, multivariate analysis of variance, and Gaussian graphical models, better results are obtained.
- Research Article
98
- 10.2307/1938927
- Feb 1, 1991
- Ecology
We examined the ability of eigenvalue tests to distinguish field—collected from random, assemblage structure data sets. Eight published time series of species abundances were used in the analysis, including data sets for: fishes, birds, mammals, stream benthos, and crabs. To test the efficacy of eigenvalue tests, we constructed 1000 randomly generated data sets for each real data set, whose means and variances were identical to the means and variances of the original data matrices. The data sets were then subjected to a principal components analysis (PCA) and eigenvalue tests used to identify significant eigenvalues for both correlation and covariance matrix solutions. We also examined the effects of: (1) number of species (= number of variables), (2) number of samples (= replication), (3) variance structure, on the performance of the test. Using PCA's based on the correlation matrix and with sample sizes typically encountered in the field, the eigenvalue tests generally performed at the .05 level when a = .01. Slightly poorer results were obtained with the covariance matrix. Increasing the number of samples to at least three times the number of species generally gave a level coverage of an a level test (i.e., a = .05, .01). Increasing variance in the data set only affected test outcomes at levels of replication less than twice the number of species. We conclude that the eigenvalue tests can be used to detect patterns in PCA's of assemblage structure data, if the number of samples is at least three times the number of species and either a covariance or correlation matrix solution is used. It is assumed that these patterns represent ecologically meaningful patterns of variation.
- Book Chapter
1
- 10.5772/8949
- Apr 1, 2010
Principal component analysis (PCA), which is also known as Karhunen-Loeve (KL) transform, is a classical statistic technique that has been applied to many fields, such as knowledge representation, pattern recognition and image compression. The objective of PCA is to reduce the dimensionality of dataset and identify new meaningful underlying variables. The key idea is to project the objects to an orthogonal subspace for their compact representations. It usually involves a mathematical procedure that transforms a number of correlated variables into a smaller number of uncorrelated variables, which are called principal components. The first principal component accounts for as much of the variability in the dataset as possible, and each succeeding component accounts for as much of the remaining variability as possible. In pattern recognition, PCA technique was first applied to the representation of human face images by Sirovich and Kirby in [1,2]. This then led to the well-known Eigenfaces method for face recognition proposed by Turk and Penland in [3]. Since then, there has been an extensive literature that addresses both the theoretical aspect of the Eigenfaces method and its application aspect [4-6]. In image compression, PCA technique has also been widely applied to the remote hyperspectral imagery for classification and compression [7,8]. Nevertheless, it can be noted that in the classical 1DPCA scheme the 2D data sample (e.g. image) must be initially converted to a 1D vector form. The resulting sample vector will lead to a high dimensional vector space. It is consequently difficult to evaluate the covariance matrix accurately when the sample vector is very long and the number of training samples is small. Furthermore, it can also be noted that the projection of a sample on each principal orthogonal vector is a scale. Obviously, this will cause the sample data to be over-compressed. In order to solve this kind of dimensionality problem, Yang et al. [9,10] proposed the 2D-PCA approach. The basic idea is to directly use a set of matrices to construct the corresponding covariance matrix instead of a set of vectors. Compared with the covariance matrix of 1D-PCA, one can note that the size of the covariance matrix using 2D-PCA is much smaller. This improves the computational efficiency. Furthermore, it can be noted that the projection of a sample on each principal orthogonal vector is a vector. Thus, the problem of over-compression is alleviated in the 2DPCA scheme. In addition, Wang et al. [11] proposed that the 2D-PCA was equivalent to a special case of the block-based PCA, and emphasized that this kind of block-based methods had been used for face recognition in a number of systems. 2
- Book Chapter
1
- 10.1007/978-90-481-8776-8_34
- Jan 1, 2010
Principal Component Analysis (PCA) is a technique to transform the original set of variables into a smaller set of linear combinations that account for most of the original set variance. The data reduction based on the classical PCA is fruitless if outlier is present in the data. The decomposed classical covariance matrix is very sensitive to outlying observations. ROBPCA is an effective PCA method combining two advantages of both projection pursuit and robust covariance estimation. The estimation is computed with the idea of minimum covariance determinant (MCD) of covariance matrix. The limitation of MCD is when covariance determinant almost equal zero. This paper proposes PCA using the minimum vector variance (MVV) as new measure of robust PCA to enhance the result. MVV is defined as a minimization of sum of square length of the diagonal of a parallelotope to determine the location estimator and covariance matrix. The usefulness of MVV is not limited to small or low dimension data set and to non-singular or singular covariance matrix. The MVV algorithm, compared with FMCD algorithm, has a lower computational complexity; the complexity of VV is of order O(p 2).
- Research Article
3
- 10.1007/s00357-015-9184-0
- Oct 1, 2015
- Journal of Classification
A model is presented for analyzing general multivariate data. The model puts as its prime objective the dimensionality reduction of the multivariate problem. The only requirement of the model is that the input data to the statistical analysis be a covariance matrix, a correlation matrix, or more generally a positive semi-definite matrix. The model is parameterized by a scale parameter and a shape parameter both of which take on non-negative values smaller than unity. We first prove a wellknown heuristic for minimizing rank and establish the conditions under which rank can be replaced with trace. This result allows us to solve our rank minimization problem as a Semi-Definite Programming (SDP) problem by a number of available solvers. We then apply the model to four case studies dealing with four well-known problems in multivariate analysis. The first problem is to determine the number of underlying factors in factor analysis (FA) or the number of retained components in principal component analysis (PCA). It is shown that our model determines the number of factors or components more efficiently than the commonly used methods. The second example deals with a problem that has received much attention in recent years due to its wide applications, and it concerns sparse principal components and variable selection in PCA. When applied to a data set known in the literature as the pitprop data, we see that our approach yields PCs with larger variances than PCs derived from other approaches. The third problem concerns sensitivity analysis of the multivariate models, a topic not widely researched in the sequel due to its difficulty. Finally, we apply the model to a difficult problem in PCA known as lack of scale invariance in the solutions of PCA. This is the problem that the solutions derived from analyzing the covariance matrix in PCA are generally different (and not linearly related to) the solutions derived from analyzing the correlation matrix. Using our model, we obtain the same solution whether we analyze the correlation matrix or the covariance matrix since the analysis utilizes only the signs of the correlations/covariances but not their values. This is where we introduce a new type of PCA, called Sign PCA, which we speculate on its applications in social sciences and other fields of science.
- Research Article
3
- 10.15866/iree.v8i1.1738
- Feb 28, 2013
- International Review of Electrical Engineering-iree
In this study, the development of nonnegative matrix factorization aided principal component analysis (NMF-PCA) algorithm is proposed to solve the problem that the covariance matrix cannot be computed due to the extremely high vector space caused by the “matrix-to-vector” transformation when principal component analysis (PCA) is applied to high-resolution image compression, which is further employed for PD gray image compression and recognition. In the proposed NMF-PCA algorithm, nonnegative matrix factorization (NMF) is firstly employed to decompose the high-resolution image into base matrices W and coefficient matrices H with lower dimension. Then, PCA is adopted to extract several principal components from the vectors of W and H as features. A fuzzy C-means (FCM) clustering method is responsible for PD classification and features evaluation. Using a traditional pulse current detector for PD experiment, 177 gray images associated with four PD defect types are obtained for NMF-PCA testing. The results of algorithm performance evaluation show that NMF and PCA both have fast responding time when r is less 3, which is suitable for PD on-line analysis and diagnosis. The recognition results of experimental PD samples demonstrate that only is the feature set FH extracted from the coefficient matrix H fit for PCA compression of PD gray images. Meanwhile, the maximum successful clustering rate 0.9661 is achieved by 3D features of FH with r = 2, which is much higher than 0.8023 of traditional PRPD operators. In addition, the FCM validity measures report that the features obtained by NMF-PCA have better aggregation characteristic than PRPD statistical operators. The obtained results demonstrate that the proposed NMF-PCA algorithm could provide an effective tool for PD diagnosis, and it is easy to extend to other image or matrix applications
- Conference Article
5
- 10.1109/icig.2009.47
- Sep 1, 2009
In this paper, a novel wood identification approach based on the two-directional two-dimensional PCA ((2D)2PCA) method is proposed in contrast to the principal component analysis(PCA), two-dimensional PCA(2DPCA), column-directional 2DPCA(c2DPCA). PCA is a classical technique used to find patterns in high dimensional data. Wood identification based on PCA must transform 2D image matrix into 1D vector, and then calculate principal components from these 1D vectors. 2DPCA, c2DPCA and (2D)2PCA methods are based on 2D image matrix as opposed to classical PCA. All these wood identification methods involve seven identical steps: (1) calculating sample’s mean; (2) demeaning all images; (3) calculating the demeaned sample’s covariance matrix; (4) eigenvalues and eigenvectors decomposition of covariance matrix; (5) eigenvectors selection according to the largest eigenvalues to construct feature space; (6) extracting features by projecting image onto feature space; (7) classifying by the nearest neighbor classifier with Euclidean distance in feature space. Experiments using PCA, 2DPCA, c2DPCA, (2D)2PCA methods are involved in this paper. By comparing PCA, 2DPCA and c2DPCA in these experiments, it is revealed that (2D)2PCA is a more efficient method in wood identification.
- Research Article
- 10.1109/jrproc.1962.288291
- May 1, 1962
- Proceedings of the IRE
Control engineering, erected on the foundations of feedback theory and linear system analysis, brings together the fundamental concepts of communication theory, network theory, nonlinear mechanics, optimization theory, and statistical design theory. These various elements are combined in the modern research on adaptive control systems-research which provides the basis from which control theory contributes to system engineering and to the understanding of the behavioral, social, and biological sciences.
- Research Article
8
- 10.1103/physreve.88.062127
- Dec 16, 2013
- Physical Review E
We design nonlinear functions for the transmission of a small signal with non-Gaussian noise and perform experiments to characterize their responses. Using statistical design theory [A. Ichiki and Y. Tadokoro, Phys. Rev. E 87, 012124 (2013)], a static nonlinear function is estimated from the probability density function of the given noise in order to maximize the signal-to-noise ratio of the output. Using an electronic system that implements the optimized nonlinear function, we confirm the recovery of a small signal from a signal with non-Gaussian noise. In our experiment, the non-Gaussian noise is a mixture of Gaussian noises. A similar technique is also applied to the optimization of the threshold value of the function. We find that, for non-Gaussian noise, the response of the optimized nonlinear systems is better than that of the linear system.