Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
18 result(s) for "Xu, Yunpei"
Sort by:
scCAD: Cluster decomposition-based anomaly detection for rare cell identification in single-cell expression data
Single-cell RNA sequencing (scRNA-seq) technologies have become essential tools for characterizing cellular landscapes within complex tissues. Large-scale single-cell transcriptomics holds great potential for identifying rare cell types critical to the pathogenesis of diseases and biological processes. Existing methods for identifying rare cell types often rely on one-time clustering using partial or global gene expression. However, these rare cell types may be overlooked during the clustering phase, posing challenges for their accurate identification. In this paper, we propose a Cluster decomposition-based Anomaly Detection method (scCAD), which iteratively decomposes clusters based on the most differential signals in each cluster to effectively separate rare cell types and achieve accurate identification. We benchmark scCAD on 25 real-world scRNA-seq datasets, demonstrating its superior performance compared to 10 state-of-the-art methods. In-depth case studies across diverse datasets, including mouse airway, brain, intestine, human pancreas, immunology data, and clear cell renal cell carcinoma, showcase scCAD’s efficiency in identifying rare cell types in complex biological scenarios. Furthermore, scCAD can correct the annotation of rare cell types and identify immune cell subtypes associated with disease, thereby offering valuable insights into disease progression. Identifying rare cells is essential for advancing our understanding of complex biological systems and disease mechanisms. Here, authors propose scCAD, a method that combines cluster decomposition and anomaly detection to effectively identify rare cell types across diverse biological scenarios.
Single-Cell Clustering Based on Shared Nearest Neighbor and Graph Partitioning
Clustering of single-cell RNA sequencing (scRNA-seq) data enables discovering cell subtypes, which is helpful for understanding and analyzing the processes of diseases. Determining the weight of edges is an essential component in graph-based clustering methods. While several graph-based clustering algorithms for scRNA-seq data have been proposed, they are generally based on k-nearest neighbor (KNN) and shared nearest neighbor (SNN) without considering the structure information of graph. Here, to improve the clustering accuracy, we present a novel method for single-cell clustering, called structural shared nearest neighbor-Louvain (SSNN-Louvain), which integrates the structure information of graph and module detection. In SSNN-Louvain, based on the distance between a node and its shared nearest neighbors, the weight of edge is defined by introducing the ratio of the number of the shared nearest neighbors to that of nearest neighbors, thus integrating structure information of the graph. Then, a modified Louvain community detection algorithm is proposed and applied to identify modules in the graph. Essentially, each community represents a subtype of cells. It is worth mentioning that our proposed method integrates the advantages of both SNN graph and community detection without the need for tuning any additional parameter other than the number of neighbors. To test the performance of SSNN-Louvain, we compare it to five existing methods on 16 real datasets, including nonnegative matrix factorization, single-cell interpretation via multi-kernel learning, SNN-Cliq, Seurat and PhenoGraph. The experimental results show that our approach achieves the best average performance in these datasets.
SAGE: Spatially Aware Gene Selection and Dual‐View Embedding Fusion for Domain Identification in Spatial Transcriptomics
Despite enabling high‐resolution mapping of gene expression within tissues, spatial transcriptomics (ST) still faces challenges in accurately segmenting spatial domains due to complex tissue architecture and limitations of current methods. Most approaches rely on local spatial priors, lack gene‐level interpretability, and fall short in capturing structure‐discriminative genes or long‐range functional relationships, limiting their ability to resolve biologically meaningful architectures. We present Spatially Aware Gene selection and dual‐view Embedding fusion (SAGE), a unified and reproducible framework for domain identification in spatial transcriptomics that combines topic‐driven gene selection with dual‐view embedding fusion to address these gaps. SAGE integrates non‐negative matrix factorization (NMF)‐based topic modeling with classifier‐based importance scoring to identify highly spatially informative genes, and fuses a local expression graph with a topic‐driven non‐local graph via consensus refinement and contrastive graph representation learning to jointly learn spatial and functional embeddings. Evaluated on 34 real‐world datasets, SAGE not only outperforms existing methods in clustering accuracy but also reveals functionally coherent regions and interpretable gene expression patterns. In case studies, SAGE reveals spatial heterogeneity associated with a pre‐malignant activation state in human breast cancer. Moreover, in zebrafish melanoma, it refines the tumor–muscle interface into transcriptionally distinct subdomains and uncovers shared vascular signatures between anatomically separate tissues. Together, these results demonstrate that SAGE can be used not only for accurate spatial domain delineation across diverse ST platforms, but also for dissecting microenvironmental niches and long‐range tissue interactions underlying disease progression. SAGE is a unified framework for spatial domain identification in spatial transcriptomics that jointly models tissue architecture and gene programs. Topic‐driven gene selection (NMF plus classifier‐based scoring) highlights spatially informative genes, while dual‐view graph embedding fuses local expression and non‐local functional relations. Across 34 datasets, SAGE improves clustering and yields interpretable domains, revealing microenvironmental niches and long‐range tissue interactions in cancer and melanoma.
A Hybrid Clustering Algorithm for Identifying Cell Types from Single-Cell RNA-Seq Data
Single-cell RNA sequencing (scRNA-seq) has recently brought new insight into cell differentiation processes and functional variation in cell subtypes from homogeneous cell populations. A lack of prior knowledge makes unsupervised machine learning methods, such as clustering, suitable for analyzing scRNA-seq. However, there are several limitations to overcome, including high dimensionality, clustering result instability, and parameter adjustment complexity. In this study, we propose a method by combining structure entropy and k nearest neighbor to identify cell subpopulations in scRNA-seq data. In contrast to existing clustering methods for identifying cell subtypes, minimized structure entropy results in natural communities without specifying the number of clusters. To investigate the performance of our model, we applied it to eight scRNA-seq datasets and compared our method with three existing methods (nonnegative matrix factorization, single-cell interpretation via multikernel learning, and structural entropy minimization principle). The experimental results showed that our approach achieves, on average, better performance in these datasets compared to the benchmark methods.
Inference and visualization of multi-scale cell tree for decoding functional diversity with scMustree
Deciphering cellular heterogeneity and functional diversity is crucial to understanding biological systems and disease mechanisms. Although single-cell RNA sequencing (scRNA-seq) has revolutionized this field, existing methods often fail to characterize the functional associations between cell populations due to the discrete nature of clustering results. We introduce scMustree, a tree- construction algorithm that enables in-depth exploration of functional associations and transitional states among cellular populations. scMustree combines top-down iterative decomposition for high- purity leaf nodes with bottom-up merging based on quantitative functional distances, preserving local continuity and global structural clarity. Benchmarking on real datasets shows that scMustree not only outperforms existing methods in clustering accuracy but also reveals functionally coherent subtypes, anatomically related cell types, and developmentally connected transitions. In case studies, it identifies Alzheimer’s disease-associated microglial subtypes and captures spatial cell distribution patterns of cell types in developed embryos, demonstrating its utility as a powerful tool for multi-scale exploration of cellular heterogeneity.
Cluster decomposition-based anomaly detection for rare cell identification in single-cell expression data
Single-cell RNA sequencing (scRNA-seq) technologies have been widely used to characterize cellular landscapes in complex tissues. Large-scale single-cell transcriptomics holds great potential for identifying rare cell types critical to the pathogenesis of diseases and biological processes. Existing methods for identifying rare cell types often rely on one-time clustering using partial or global gene expression. However, these rare cell types may be overlooked in the initial clustering step, making them difficult to distinguish. In this paper, we propose a Cluster decomposition-based Anomaly Detection method (scCAD), which iteratively decomposes clusters based on the most differential signals in each cluster to effectively separate rare cell types and achieve accurate identification. We benchmark scCAD on 25 real-world scRNA-seq datasets, demonstrating its superior performance compared to 10 state-of-the-art methods. In-depth case studies across diverse datasets, including mouse airway, brain, intestine, human pancreas, immunology data, and clear cell renal cell carcinoma, showcase scCAD's efficiency in identifying rare cell types in complex biological scenarios. Furthermore, scCAD can correct the annotation of rare cell types and identify immune cell subtypes associated with disease, providing new insights into disease progression.Competing Interest StatementThe authors have declared no competing interest.Footnotes* https://www.ncbi.nlm.nih.gov/geo/* https://www.ebi.ac.uk/arrayexpress/* https://www.ncbi.nlm.nih.gov/sra* https://www.10xgenomics.com/* http://atlas.gs.washington.edu/worm-rna/docs/* https://portals.broadinstitute.org/single_cell* https://github.com/OSU-BMBL/marsgt
TopoLa: A Universal Framework to Enhance Cell Representations for Single-cell and Spatial Omics through Topology-encoded Latent Hyperbolic Geometry
Recent advances in cellular research demonstrate that scRNA-seq characterizes cellular heterogeneity, while spatial transcriptomics reveals the spatial distribution of gene expression. Cell representation is the fundamental issue in the two fields. Here, we propose Topology-encoded Latent Hyperbolic Geometry (TopoLa), a computational framework enhancing cell representations by capturing fine-grained intercellular topological relationships. The framework introduces a new metric, TopoLa distance (TLd), which quantifies the geometric distance between cells within latent hyperbolic space, capturing the network’s topological structure more effectively. With this framework, the cell representation can be enhanced considerably by performing convolution on its neighboring cells. Performance evaluation across seven biological tasks, including scRNA-seq data clustering and spatial transcriptomics domain identification, shows that TopoLa significantly improves the performance of several state-of-the-art models. These results underscore the generalizability and robustness of TopoLa, establishing it as a valuable tool for advancing both biological discovery and computational methodologies.
TopoLa: A Universal Framework to Enhance Cell Representations for Single-cell and Spatial Omics through Topology-encoded Latent Hyperbolic Geometry
Recent advances in cellular research demonstrate that scRNA-seq characterizes cellular heterogeneity, while spatial transcriptomics reveals the spatial distribution of gene expression. Cell representation is the fundamental issue in the two fields. Here, we propose Topology-encoded Latent Hyperbolic Geometry (TopoLa), a computational framework enhancing cell representations by capturing fine-grained intercellular topological relationships. The framework introduces a new metric, TopoLa distance (TLd), which quantifies the geometric distance between cells within latent hyperbolic space, capturing the network's topological structure more effectively. With this framework, the cell representation can be enhanced considerably by performing convolution on its neighboring cells. Performance evaluation across seven biological tasks, including scRNA-seq data clustering and spatial transcriptomics domain identification, shows that TopoLa significantly improves the performance of several state-of-the-art models. These results underscore the generalizability and robustness of TopoLa, establishing it as a valuable tool for advancing both biological discovery and computational methodologies.
ClusterMine: a Knowledge-integrated Clustering Approach based on Expression Profiles of Gene Sets
Clustering analysis is essential for understanding complex biological data. In widely used methods such as hierarchical clustering (HC) and consensus clustering (CC), expression profiles of all genes are often used to assess similarity between samples for clustering. These methods output sample clusters, but are not able to provide information about which gene sets (functions) contribute most to the clustering. So interpretability of their results is limited. We hypothesized that integrating prior knowledge of annotated biological processes would not only achieve satisfying clustering performance but also, more importantly, enable potential biological interpretation of clusters. Here we report ClusterMine, a novel approach that identifies clusters by assessing functional similarity between samples through integrating known annotated gene sets, e.g., in Gene Ontology. In addition to outputting cluster membership of each sample as conventional approaches do, it outputs gene sets that are most likely to contribute to the clustering, a feature facilitating biological interpretation. Using three cancer datasets, two single cell RNA-sequencing based cell differentiation datasets, one cell cycle dataset and two datasets of cells of different tissue origins, we found that ClusterMine achieved similar or better clustering performance and that top-scored gene sets prioritized by ClusterMine are biologically relevant. ClusterMine is implemented as an R package and is freely available at: www.genemine.org/clustermine.php
Experimental Study on Similarity Simulation of Mechanical Properties of Coal Rock Mass in Folded Structural Zones
To thoroughly investigate the mechanisms behind coal and gas outbursts in folded structural areas, we conducted similarity simulation experiments using a custom-built apparatus designed to replicate these structures. The objective was to analyze the stress distribution characteristics of coal rock masses under horizontal structural stress within folded zones. The experimental outcomes reveal that, under horizontal loading, shear cracks progressively develop along layer directions within the anticline wing, anticline axis, and syncline axis, evolving continuously along the interlayer direction. In these folded structures, horizontal stress consistently remains compressive, with the highest compressive stress concentrations observed at the anticline axis, followed by the wings and turning points of the anticline, and the lowest in the syncline axis area. The stress coefficient (k) in the anticline axis reached values as high as 3.18, while the syncline axis exhibited much lower stress concentrations, with k values of 0.66. Vertically, the anticline axis and its wings primarily experience tensile stress, whereas the syncline and its wings mainly undergo vertical compressive stress. The anticline axis region, subjected to horizontal structural stress, tends to develop tension cracks, which adversely affect gas retention. The combination of horizontal tension and vertical tensile stress in this region reduces the risk of coal and gas outbursts. Conversely, the syncline axis area, experiencing triaxial compressive stress, exhibits a higher degree of stress concentration and superior gas sealing capacity, rendering it more vulnerable to coal and gas outbursts. These findings provide essential insights for refining coal mining methodologies in fold structures, particularly for addressing the safety challenges posed by coal and gas outbursts.