Living on the edge: phase transitions in convex programs with random data
Recent research indicates that many convex optimization problems with random constraints exhibit a phase transition as the number of constraints increases. For example, this phenomenon emerges in the l_1 minimization method for identifying a sparse vector from random linear measurements. Indeed, the l_1 approach succeeds with high probability when the number of measurements exceeds a threshold that depends on the sparsity level; otherwise, it fails with high probability. This paper provides the first rigorous analysis that explains why phase transitions are ubiquitous in random convex optimization problems. It also describes tools for making reliable predictions about the quantitative aspects of the transition, including the location and the width of the transition region. These techniques apply to regularized linear inverse problems with random measurements, to demixing problems under a random incoherence model, and also to cone programs with random affine constraints. The applied results depend on foundational research in conic geometry. This paper introduces a summary parameter, called the statistical dimension, that canonically extends the dimension of a linear subspace to the class of convex cones. The main technical result demonstrates that the sequence of intrinsic volumes of a convex cone concentrates sharply around the statistical dimension. This fact leads to accurate bounds on the probability that a randomly rotated cone shares a ray with a fixed cone.
- Research Article
88
- 10.1007/s10107-012-0540-0
- May 1, 2012
- Mathematical Programming
This note presents a unified analysis of the recovery of simple objects from random linear measurements. When the linear functionals are Gaussian, we show that an s-sparse vector in \({\mathbb{R}^n}\) can be efficiently recovered from 2s log n measurements with high probability and a rank r, n × n matrix can be efficiently recovered from r(6n − 5r) measurements with high probability. For sparse vectors, this is within an additive factor of the best known nonasymptotic bounds. For low-rank matrices, this matches the best known bounds. We present a parallel analysis for block-sparse vectors obtaining similarly tight bounds. In the case of sparse and block-sparse signals, we additionally demonstrate that our bounds are only slightly weakened when the measurement map is a random sign matrix. Our results are based on analyzing a particular dual point which certifies optimality conditions of the respective convex programming problem. Our calculations rely only on standard large deviation inequalities and our analysis is self-contained.
- Research Article
482
- 10.1007/s10208-007-9011-z
- Dec 18, 2007
- Foundations of Computational Mathematics
We propose a new approach for nonadaptive dimensionality reduction of manifold-modeled data, demonstrating that a small number of random linear projections can preserve key information about a manifold-modeled signal. We center our analysis on the effect of a random linear projection operator Φ:ℝN→ℝM, M<N, on a smooth well-conditioned K-dimensional submanifold ℳ⊂ℝN. As our main theoretical contribution, we establish a sufficient number M of random projections to guarantee that, with high probability, all pairwise Euclidean and geodesic distances between points on ℳ are well preserved under the mapping Φ. Our results bear strong resemblance to the emerging theory of Compressed Sensing (CS), in which sparse signals can be recovered from small numbers of random linear measurements. As in CS, the random measurements we propose can be used to recover the original data in ℝN. Moreover, like the fundamental bound in CS, our requisite M is linear in the “information level” K and logarithmic in the ambient dimension N; we also identify a logarithmic dependence on the volume and conditioning of the manifold. In addition to recovering faithful approximations to manifold-modeled signals, however, the random projections we propose can also be used to discern key properties about the manifold. We discuss connections and contrasts with existing techniques in manifold learning, a setting where dimensionality reducing mappings are typically nonlinear and constructed adaptively from a set of sampled training data.
- Conference Article
9
- 10.23919/ecc.2013.6669266
- Jul 1, 2013
We consider convex optimization problems with N randomly drawn convex constraints. Previous work has shown that the tails of the distribution of the probability that the optimal solution subject to these constraints will violate the next random constraint, can be bounded by a binomial distribution. In this paper we extend these results to the violation probability of convex combinations of optimal solutions of optimization problems with random constraints and different cost objectives. This extension has interesting applications to distributed multi-agent consensus algorithms in which the decision vectors of the agents are subject to random constraints and the agents' goal is to achieve consensus on a common value of the decision vector that satisfies the constraints. We give explicit bounds on the tails of the probability that the agents' decision vectors at an arbitrary iteration of the consensus protocol violate further constraint realizations. In a numerical experiment we apply these results to a model predictive control problem in which the agents aim to achieve consensus on a control sequence subject to random terminal constraints.
- Research Article
4
- 10.1007/s00034-017-0730-3
- Dec 22, 2017
- Circuits, Systems, and Signal Processing
The success of orthogonal matching pursuit (OMP) in the sparse signal recovery heavily depends on its ability for correct support recovery. Based on a support recovery guarantee for OMP expressed in terms of the mutual coherence, and a result about the concentration of the extreme singular values of a Gaussian random matrix, this paper proposes a preconditioning method for increasing the recovery rate of OMP from random and noisy measurements. Compared to several existing preconditionings, the proposed method can reduce the mutual coherence with a proven high probability. Simultaneously, the proposed preconditioning can also succeed with a high probability in providing slight signal-to-noise ratio reduction, which is empirically shown to be less severe than that caused by a recently suggested technique for the noisy case. The simulations show the advantages of the proposed preconditioning over other currently relevant ones in terms of both the performance improvement for OMP, and computation time.
- Research Article
- 10.1093/imaiai/iaag006
- Feb 19, 2026
- Information and Inference: A Journal of the IMA
Understanding the stochastic behavior of random projections of geometric sets constitutes a fundamental problem in high dimension probability that finds wide applications in diverse fields. This paper provides a kinematic description for the behavior of Gaussian random projections of closed convex cones, in analogy to that of randomly rotated cones studied in Amelunxen et al. (2014, Inf. Inference J. IMA, 3, 224–294). Formally, let $K$ be a closed convex cone in $\mathbb{R}^{n}$, and $G\in \mathbb{R}^{m\times n}$ be a Gaussian matrix with i.i.d. $\mathscr{N}(0,1)$ entries. We show that $GK\equiv \{G\mu : \mu \in K\}$ behaves like a randomly rotated cone in $\mathbb{R}^{m}$ with statistical dimension $\min \{\delta (K),m\}$, in the following kinematic sense: for any fixed closed convex cone $L$ in $\mathbb{R}^{m}$,$$ \begin{align*} &\delta(L)+\delta(K)\ll m\, \Rightarrow\, L\cap GK = \{0\} \hbox{ with high probability},\nonumber\\ &\delta(L)+\delta(K)\gg m\, \Rightarrow\, L\cap GK \neq \{0\} \hbox{ with high probability}. \end{align*} $$ Similar kinematic descriptions are obtained for Gaussian random pre-images, and certain Gaussian random projections of general closed convex sets. The practical utility and broad applicability of the prescribed approximate kinematic formulae are demonstrated in a number of distinct problems arising from statistical learning, mathematical programming and asymptotic geometric analysis. In particular, we prove (i) new phase transitions of the existence of cone-constrained maximum likelihood estimators in logistic regression, (ii) new phase transitions of the cost optimum of deterministic conic programs with random constraints and (iii) a local version of the Gaussian Dvoretzky-Milman theorem that describes almost deterministic, low-dimensional behaviours of subspace sections of randomly projected convex sets. The proofs of our results exploit the full strength of comparison inequalities for Gaussian processes. Compared to the conic integral geometry method in Amelunxen et al. (2014, Inf. Inference J. IMA, 3, 224–294), our method has the advantage of circumventing the rigid requirement of exact kinematic formulae that are typically unavailable for random projections and general closed convex sets.
- Research Article
51
- 10.1109/tsp.2018.2827326
- Oct 10, 2017
- IEEE Transactions on Signal Processing
Efficient estimation of wideband spectrum is of great importance for applications such as cognitive radio. Recently, sub-Nyquist sampling schemes based on compressed sensing have been proposed to greatly reduce the sampling rate. However, the important issue of quantization has not been fully addressed, particularly for high-resolution spectrum and parameter estimation. In this paper, we aim to recover spectrally-sparse signals and the corresponding parameters, such as frequency and amplitudes, from heavy quantizations of their noisy complex-valued random linear measurements, e.g. only the quadrant information. We first characterize the Cramer-Rao bound under Gaussian noise, which highlights the trade-off between sample complexity and bit depth under different signal-to-noise ratios for a fixed budget of bits. Next, we propose a new algorithm based on atomic norm soft thresholding for signal recovery, which is equivalent to proximal mapping of properly designed surrogate signals with respect to the atomic norm that motivates spectral sparsity. The proposed algorithm can be applied to both the single measurement vector case, as well as the multiple measurement vector case. It is shown that under the Gaussian measurement model, the spectral signals can be reconstructed accurately with high probability, as soon as the number of quantized measurements exceeds the order of K log n, where K is the level of spectral sparsity and $n$ is the signal dimension. Finally, numerical simulations are provided to validate the proposed approaches.
- Conference Article
6
- 10.1109/sampta.2017.8024376
- Jul 1, 2017
Recent work has demonstrated the effectiveness of gradient descent for recovering low-rank matrices from random linear measurements in a globally convergent manner. However, their performance is highly sensitive in the presence of outliers that may take arbitrary values, which is common in practice. In this paper, we propose a truncated gradient descent algorithm to improve the robustness against outliers, where the truncation is performed to rule out the contributions from samples that deviate significantly from the sample median. A restricted isometry property regarding the sample median is introduced to provide a theoretical footing of the proposed algorithm for the Gaussian orthogonal ensemble. Extensive numerical experiments are provided to validate the superior performance of the proposed algorithm.
- Conference Article
49
- 10.1109/icassp.2014.6854351
- May 1, 2014
In this paper, we analyze the security of compressed sensing (CS) as a cryptosystem. We demonstrate that random linear measurements acquired using a Gaussian i.i.d. matrix reveal only the energy of the sensed signal, and that only the energy of the measurements leaks information about the signal. We provide useful bounds for assessing the information leakage about the energy, linking those bounds to the minimum mean square error achievable by practical estimators. Moreover, we propose a simple strategy based on the normalization of the measurements which achieves, at least in theory, perfect secrecy, enabling the use of CS-based encryption in practical cryptosystems.
- Research Article
13
- 10.1073/pnas.1705490115
- Jun 25, 2018
- Proceedings of the National Academy of Sciences of the United States of America
In matrix recovery from random linear measurements, one is interested in recovering an unknown M-by-N matrix [Formula: see text] from [Formula: see text] measurements [Formula: see text], where each [Formula: see text] is an M-by-N measurement matrix with i.i.d. random entries, [Formula: see text] We present a matrix recovery algorithm, based on approximate message passing, which iteratively applies an optimal singular-value shrinker-a nonconvex nonlinearity tailored specifically for matrix estimation. Our algorithm typically converges exponentially fast, offering a significant speedup over previously suggested matrix recovery algorithms, such as iterative solvers for nuclear norm minimization (NNM). It is well known that there is a recovery tradeoff between the information content of the object [Formula: see text] to be recovered (specifically, its matrix rank r) and the number of linear measurements n from which recovery is to be attempted. The precise tradeoff between r and n, beyond which recovery by a given algorithm becomes possible, traces the so-called phase transition curve of that algorithm in the [Formula: see text] plane. The phase transition curve of our algorithm is noticeably better than that of NNM. Interestingly, it is close to the information-theoretic lower bound for the minimal number of measurements needed for matrix recovery, making it not only state of the art in terms of convergence rate, but also near optimal in terms of the matrices it successfully recovers.
- Supplementary Content
15
- 10.7907/156s-ez89.
- Jan 1, 2013
Demixing is the task of identifying multiple signals given only their sum and prior information about their structures. Examples of demixing problems include (i) separating a signal that is sparse with respect to one basis from a signal that is sparse with respect to a second basis; (ii) decomposing an observed matrix into low-rank and sparse components; and (iii) identifying a binary codeword with impulsive corruptions. This thesis describes and analyzes a convex optimization framework for solving an array of demixing problems. Our framework includes a random orientation model for the constituent signals that ensures the structures are incoherent. This work introduces a summary parameter, the statistical dimension, that reflects the intrinsic complexity of a signal. The main result indicates that the difficulty of demixing under this random model depends only on the total complexity of the constituent signals involved: demixing succeeds with high probability when the sum of the complexities is less than the ambient dimension; otherwise, it fails with high probability. The fact that a phase transition between success and failure occurs in demixing is a consequence of a new inequality in conic integral geometry. Roughly speaking, this inequality asserts that a convex cone behaves like a subspace whose dimension is equal to the statistical dimension of the cone. When combined with a geometric optimality condition for demixing, this inequality provides precise quantitative information about the phase transition, including the location and width of the transition region.
- Research Article
10224
- 10.1109/tit.2007.909108
- Dec 1, 2007
- IEEE Transactions on Information Theory
<para xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> This paper demonstrates theoretically and empirically that a greedy algorithm called Orthogonal Matching Pursuit (OMP) can reliably recover a signal with <emphasis><formula formulatype="inline"><tex>$m$</tex></formula></emphasis> nonzero entries in dimension <emphasis><formula formulatype="inline"><tex>$d$</tex> </formula></emphasis> given <emphasis><formula formulatype="inline"><tex>$ {\rm O}(m \ln d)$</tex></formula></emphasis> random linear measurements of that signal. This is a massive improvement over previous results, which require <emphasis><formula formulatype="inline"><tex>${\rm O}(m^{2})$</tex></formula></emphasis> measurements. The new results for OMP are comparable with recent results for another approach called Basis Pursuit (BP). In some settings, the OMP algorithm is faster and easier to implement, so it is an attractive alternative to BP for signal recovery problems. </para>
- Conference Article
70
- 10.1109/icassp.2007.366822
- Apr 1, 2007
Encouraging recent results in compressed sensing or compressive sampling suggest that a set of inner products with random measurement vectors forms a good representation of a source vector that is known to be sparse in some fixed basis. With quantization of these inner products, the encoding can be considered universal for sparse signals with known sparsity level. We analyze the operational rate-distortion performance of such source coding both with genie-aided knowledge of the sparsity pattern and maximum likelihood estimation of the sparsity pattern. We show that random measurements induce an additive logarithmic rate penalty, i.e., at high rates the performance with rate R + O(log R) and random measurements is equal to the performance with rate R and deterministic measurements matched to the source.
- Conference Article
- 10.1117/12.2306389
- May 24, 2018
Based on the objects’ sparse features, the compressive sensing imaging system has the unique advantage of breaking the Nyquist sampling theorem, and the target image can be reconstructed from very few random coded observations. The system is characterized by simple coding and complex decoding. It is difficult to meet the increasing real-time requirements in application because of the large time consumption by the iterative optimization algorithms. Therefore, it is a powerful way to improve the efficiency by bypassing the complex reconstruction process and extracting the target information directly from the random measurement data. In this paper, based on MNIST handwritten digital character database as an example, the object recognition method from random measurements of compressive sensing camera is explored. Firstly, the training samples in the MNIST database are coded with the observation of the random Bernoulli measurement matrix. And then the K-nearest neighbor classifier is constructed on the standardized samples, the measurements in the same measurement matrix of the target sample are put in the classifier, given the target recognition results. The experimental results show that the average recognition rate is 82.8% under the sampling rate of 0.1, and the total time to process 500 images is 0.063s. In contrast, the experiment of the traditional method by first reconstructing and then recognizing is conducted, the average recognition rate is 84.3%,and the total time to process 500 images is 48.2s. The proposed method is close to the traditional strategy in recognition accuracy, but the computational efficiency has been greatly improved (765 times), with great practical value.
- Conference Article
10
- 10.1109/allerton.2009.5394539
- Sep 1, 2009
We consider the problem of recovering a low-rank matrix M from a small number of random linear measurements. A popular and useful example of this problem is matrix completion, in which the measurements reveal the values of a subset of the entries, and we wish to fill in the missing entries (this is the famous Netflix problem). When M is believed to have low rank, one would ideally try to recover M by finding the minimum-rank matrix that is consistent with the data; this is, however, problematic since this is a nonconvex problem that is, generally, intractable. Nuclear-norm minimization has been proposed as a tractable approach, and past papers have delved into the theoretical properties of nuclear-norm minimization algorithms, establishing conditions under which minimizing the nuclear norm yields the minimum rank solution. We review this spring of emerging literature and extend and refine previous theoretical results. Our focus is on providing error bounds when M is well approximated by a low-rank matrix, and when the measurements are corrupted with noise. We show that for a certain class of random linear measurements, nuclear-norm minimization provides stable recovery from a number of samples nearly at the theoretical lower limit, and enjoys order-optimal error bounds (with high probability).
- Conference Article
9
- 10.1117/12.891933
- Sep 8, 2011
- Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
Low-rank matrix recovery addresses the problem of recovering an unknown low-rank matrix from few linear measurements. Nuclear-norm minimization is a tractible approach with a recent surge of strong theoretical backing. Analagous to the theory of compressed sensing, these results have required random measurements. For example, m >= Cnr Gaussian measurements are sufficient to recover any rank-r n x n matrix with high probability. In this paper we address the theoretical question of how many measurements are needed via any method whatsoever --- tractible or not. We show that for a family of random measurement ensembles, m >= 4nr - 4r^2 measurements are sufficient to guarantee that no rank-2r matrix lies in the null space of the measurement operator with probability one. This is a necessary and sufficient condition to ensure uniform recovery of all rank-r matrices by rank minimization. Furthermore, this value of $m$ precisely matches the dimension of the manifold of all rank-2r matrices. We also prove that for a fixed rank-r matrix, m >= 2nr - r^2 + 1 random measurements are enough to guarantee recovery using rank minimization. These results give a benchmark to which we may compare the efficacy of nuclear-norm minimization.