Conceptual

High-Dimensional Statistical Inference for Large-p Small-n Data in Multivariate Statistics

When the number of measured variables p exceeds the sample size n, the sample covariance matrix becomes singular and the classical Hotelling T-squared statistic is undefined, so the whole apparatus of multivariate inference has to be rebuilt. This area develops tests for mean vectors and covariance matrices that remain valid under p>n: dimension-reduction approaches, asymptotics-driven statistics that replace the inverse covariance with traces of powers of the covariance (Bai-Saranadasa, Chen-Qin, Srivastava-Du, Park-Ayyala), and random-projection methods such as RAPTT that repeatedly project the data into a low-dimensional space where a classical test is legal and then aggregate the resulting p-values. It also covers estimation of large covariance and precision matrices under sparsity, and the corresponding inference for discrete multivariate models (multinomial, Dirichlet-multinomial, latent Dirichlet allocation) that arise when the data are counts. Applications are drawn from gene-expression genomics, metagenomic abundance tables, text mining, and social-network data, where p in the thousands and n in the tens is the normal situation rather than the exception.