Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
156 result(s) for "Zhang, Anru"
Sort by:
AN OPTIMAL STATISTICAL AND COMPUTATIONAL FRAMEWORK FOR GENERALIZED TENSOR ESTIMATION
This paper describes a flexible framework for generalized low-rank tensor estimation problems that includes many important instances arising from applications in computational imaging, genomics, and network analysis. The proposed estimator consists of finding a low-rank tensor fit to the data under generalized parametric models. To overcome the difficulty of nonconvexity in these problems, we introduce a unified approach of projected gradient descent that adapts to the underlying low-rank structure. Under mild conditions on the loss function, we establish both an upper bound on statistical error and the linear rate of computational convergence through a general deterministic analysis. Then we further consider a suite of generalized tensor estimation problems, including sub-Gaussian tensor PCA, tensor regression, and Poisson and binomial tensor PCA. We prove that the proposed algorithm achieves the minimax optimal rate of convergence in estimation error. Finally, we demonstrate the superiority of the proposed framework via extensive experiments on both simulated and real data.
TENSOR CLUSTERING WITH PLANTED STRUCTURES
This paper studies the statistical and computational limits of high-order clustering with planted structures. We focus on two clustering models, constant high-order clustering (CHC) and rank-one higher-order clustering (ROHC), and study the methods and theory for testing whether a cluster exists (detection) and identifying the support of cluster (recovery). Specifically, we identify the sharp boundaries of signal-to-noise ratio for which CHC and ROHC detection/recovery are statistically possible. We also develop the tight computational thresholds: when the signal-to-noise ratio is below these thresholds, we prove that polynomial-time algorithms cannot solve these problems under the computational hardness conjectures of hypergraphic planted clique (HPC) detection and hypergraphic planted dense subgraph (HPDS) recovery. We also propose polynomial-time tensor algorithms that achieve reliable detection and recovery when the signal-to-noise ratio is above these thresholds. Both sparsity and tensor structures yield the computational barriers in high-order tensor clustering. The interplay between them results in significant differences between high-order tensor clustering and matrix clustering in literature in aspects of statistical and computational phase transition diagrams, algorithmic approaches, hardness conjecture, and proof techniques. To our best knowledge, we are the first to give a thorough characterization of the statistical and computational trade-off for such a double computational-barrier problem. Finally, we provide evidence for the computational hardness conjectures of HPC detection (via low-degree polynomial and Metropolis methods) and HPDS recovery (via low-degree polynomial method).
INFERENCE FOR LOW-RANK TENSORS—NO NEED TO DEBIAS
In this paper, we consider the statistical inference for several low-rank tensor models. Specifically, in the Tucker low-rank tensor PCA or regression model, provided with any estimates achieving some attainable error rate, we develop the data-driven confidence regions for the singular subspace of the parameter tensor based on the asymptotic distribution of an updated estimate by two-iteration alternating minimization. The asymptotic distributions are established under some essential conditions on the signal-to-noise ratio (in PCA model) or sample size (in regression model). If the parameter tensor is further orthogonally decomposable, we develop the methods and nonasymptotic theory for inference on each individual singular vector. For the rank-one tensor PCA model, we establish the asymptotic distribution for general linear forms of principal components and confidence interval for each entry of the parameter tensor. Finally, numerical simulations are presented to corroborate our theoretical discoveries. In all of these models, we observe that different from many matrix/vector settings in existing work, debiasing is not required to establish the asymptotic distribution of estimates or to make statistical inference on low-rank tensors. In fact, due to the widely observed statistical-computational-gap for low-rank tensor estimation, one usually requires stronger conditions than the statistical (or information-theoretic) limit to ensure the computationally feasible estimation is achievable. Surprisingly, such conditions “incidentally” render a feasible low-rank tensor inference without debiasing.
HETEROSKEDASTIC PCA
A general framework for principal component analysis (PCA) in the presence of heteroskedastic noise is introduced. We propose an algorithm called HeteroPCA, which involves iteratively imputing the diagonal entries of the sample covariance matrix to remove estimation bias due to heteroskedasticity. This procedure is computationally efficient and provably optimal under the generalized spiked covariance model. A key technical step is a deterministic robust perturbation analysis on singular subspaces, which can be of independent interest. The effectiveness of the proposed algorithm is demonstrated in a suite of problems in high-dimensional statistics, including singular value decomposition (SVD) under heteroskedastic noise, Poisson PCA, and SVD for heteroskedastic and incomplete data.
Vispro improves imaging analysis for Visium spatial transcriptomics
Spatial transcriptomics enables spatially resolved gene expression analysis, but accompanying histology images are often degraded by fiducial markers and background regions, hindering interpretation. To address this, we introduce Vispro, an end-to-end automated image processing tool optimized for 10× Visium data. Vispro includes modules for fiducial marker detection, image restoration, tissue region detection, and segmentation of disconnected tissue areas. By enhancing image quality, Vispro improves the accuracy and performance of downstream analyses, including tissue and cell segmentation, image registration, gene expression imputation guided by histological context, and spatial domain detection.
NONPARAMETRIC COVARIANCE ESTIMATION FOR MIXED LONGITUDINAL STUDIES, WITH APPLICATIONS IN MIDLIFE WOMEN’S HEALTH
In mixed longitudinal studies, a group of subjects enter the study at different ages (cross-sectional) and are followed for successive years (longitudinal). In the context of such studies, we consider nonparametric covariance estimation with samples of noisy and partially observed functional trajectories. The proposed algorithm is based on a noniterative sequential-aggregation scheme with only basic matrix operations and closed-form solutions in each step. The good performance of the proposed method is supported by both theory and numerical experiments. We also apply the proposed procedure to a study on the working memory of midlife women, based on data from the Study of Women’s Health Across the Nation (SWAN).
REGRESSION ANALYSIS FOR MICROBIOME COMPOSITIONAL DATA
One important problem in microbiome analysis is to identify the bacterial taxa that are associated with a response, where the microbiome data are summarized as the composition of the bacterial taxa at different taxonomic levels. This paper considers regression analysis with such compositional data as covariates. In order to satisfy the subcompositional coherence of the results, linear models with a set of linear constraints on the regression coefficients are introduced. Such models allow regression analysis for subcompositions and include the log-contrast model for compositional covariates as a special case. A penalized estimation procedure for estimating the regression coefficients and for selecting variables under the linear constraints is developed. A method is also proposed to obtain debiased estimates of the regression coefficients that are asymptotically unbiased and have a joint asymptotic multivariate normal distribution. This provides valid confidence intervals of the regression coefficients and can be used to obtain the p-values. Simulation results show the validity of the confidence intervals and smaller variances of the debiased estimates when the linear constraints are imposed. The proposed methods are applied to a gut microbiome data set and identify four bacterial genera that are associated with the body mass index after adjusting for the total fat and caloric intakes.
Increase in antioxidant capacity associated with the successful subclone of hypervirulent carbapenem-resistant Klebsiella pneumoniae ST11-KL64
The acquisition of exogenous mobile genetic material imposes an adaptive burden on bacteria, whereas the adaptational evolution of virulence plasmids upon entry into carbapenem-resistant Klebsiella pneumoniae (CRKP) and its impact remains unclear. To better understand the virulence in CRKP, we characterize virulence plasmids utilizing a large genomic data containing 1219  K. pneumoniae from our long-term surveillance and publicly accessible databases. Phylogenetic evaluation unveils associations between distinct virulence plasmids and serotypes. The sub-lineage ST11-KL64 CRKP acquires a pK2044-like virulence plasmid from ST23-KL1 hypervirulent K. pneumoniae , with a 2698 bp region deletion in all ST11-KL64. The deletion is observed to regulate methionine metabolism, enhance antioxidant capacity, and further improve survival of hypervirulent CRKP in macrophages. The pK2044-like virulence plasmid discards certain sequences to enhance survival of ST11-KL64, thereby conferring an evolutionary advantage. This work contributes to multifaceted understanding of virulence and provides insight into potential causes behind low fitness costs observed in bacteria. Plasmid acquisition imposes an adaptive burden, which can be ameliorated by host-plasmid coevolution. Here, the authors characterise virulence plasmids of carbapenem-resistant Klebsiella pneumoniae , and show the discard of certain sequences to enhance survival, conferring an evolutionary advantage.
CROSS
The completion of tensors, or high-order arrays, attracts significant attention in recent research. Current literature on tensor completion primarily focuses on recovery from a set of uniformly randomly measured entries, and the required number of measurements to achieve recovery is not guaranteed to be optimal. In addition, the implementation of some previous methods are NP-hard. In this article, we propose a framework for lowrank tensor completion via a novel tensor measurement scheme that we name Cross. The proposed procedure is efficient and easy to implement. In particular, we show that a third-order tensor of Tucker rank-(r₁, r₂, r₃) in p₁-by-p₂-by-p₃ dimensional space can be recovered from as few as r₁r₂r₃ + r₁(p₁ − r1) + r₂(p₂ − r₂) + r₃(p₃ − r₃) noiseless measurements, which matches the sample complexity lower bound. In the case of noisy measurements, we also develop a theoretical upper bound and the matching minimax lower bound for recovery error over certain classes of low-rank tensors for the proposed procedure. The results can be further extended to fourth or higher-order tensors. Simulation studies show that the method performs well under a variety of settings. Finally, the procedure is illustrated through a real dataset in neuroimaging.
TEMPTED: time-informed dimensionality reduction for longitudinal microbiome studies
Longitudinal studies are crucial for understanding complex microbiome dynamics and their link to health. We introduce TEMPoral TEnsor Decomposition (TEMPTED), a time-informed dimensionality reduction method for high-dimensional longitudinal data that treats time as a continuous variable, effectively characterizing temporal information and handling varying temporal sampling. TEMPTED captures key microbial dynamics, facilitates beta-diversity analysis, and enhances reproducibility by transferring learned representations to new data. In simulations, it achieves 90% accuracy in phenotype classification, significantly outperforming existing methods. In real data, TEMPTED identifies vaginal microbial markers linked to term and preterm births, demonstrating robust performance across datasets and sequencing platforms.