Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
13
result(s) for
"Zhai, Zhiqian"
Sort by:
CIPHER: An end-to-end framework for designing optimized aggregated spatial transcriptomics experiments
by
Xie, Fangming
,
De Ocampo, Haley
,
Li, Jingyi Jessica
in
Algorithms
,
Animals
,
Biology and Life Sciences
2026
Most imaging-based spatial transcriptomics methods measure individual genes, which limits scalability and typically requires integration with scRNA-seq to recover full cellular states. Recent approaches such as CISI, FISHnCHIPs, and ATLAS address this limitation by measuring aggregate transcriptional signatures, where multiple genes are pooled into each channel to increase throughput. While aggregate measurements improve scalability, they shift the problem from gene selection to feature design. For effective integration with scRNA-seq, these signatures must be not only discriminative in transcriptional space but also straightforward to measure, with balanced signal, sufficient dynamic range, and robustness to experimental noise. By optimizing decoding accuracy in isolation, existing methods leave substantial performance on the table.
We present CIPHER (Cell Identity Projection using Hybridization Encoding Rules), a neural-network framework that jointly optimizes the experimental encoding matrix, i.e., the way that genes are aggregated to signatures, and the downstream cell embedding. CIPHER integrates the physical limits of imaging assays directly into its loss function, shaping the latent space to maximize discriminability while maintaining robustness to measurement noise and signal constraints. Using a large-scale mouse brain scRNA-seq reference, we show that CIPHER-designed encodings yield latent spaces with improved cell-type separability, uniform signal utilization, and greater resilience to hybridization variability, resulting in higher decoding accuracy from both simulated and experimental data.
CIPHER formulates aggregate signature design as a joint optimization problem over decoding accuracy and experimental measurability. This enables systematic, scRNA-seq-aligned feature design for scalable spatial transcriptomics based on aggregate measurements.
Code and documentation are available at https://github.com/wollmanlab/Design/.
Journal Article
BATF2 is a glutamine-responsive tumour suppressor required for type-I interferon-dependent anti-tumour immunity
2025
Recent evidence highlights the significance of a new type of tumour suppressors, which are not frequently mutated but inhibited by metabolic cues in cancers. Here, we identify BATF2 as a tumour suppressor whose expression is epigenetically silenced by glutamine in Head and Neck Squamous Cell Carcinomas (HNSCC).
BATF2
correlates with type-I interferon and Th1 signatures in human HNSCC, with correlation coefficients even stronger than those of the positive control,
STING
. The phosphorylation of BATF2 at serine 227 promotes the oligomerization of STING.
BATF2
deficiency or high glutamine levels result in higher oxygen consumption rates and metabolic profiles unfavorable for type-I interferon production. An isocaloric glutamine-rich diet abolishes STING-mediated effector cell expansion in tumours, weakening STING agonist-induced tumour control. Cancer cell-specific BATF2 expression promotes an Id2-centered T-cell effector signature, reduces T-cell exhaustion, and triggers spontaneous HNSCC rejection in a type-I interferon-dependent fashion. Utilizing syngeneic subcutaneous, orthotopic, and 24-week-long cigarette smoke carcinogen-induced HNSCC models, we demonstrate that host
Batf2
deficiency results in increased infiltration of CD206
+
myeloid cells and reduced effector CD8
+
T-cells, accelerating the initiation of cancers. Overall, we reveal a tumour suppressor
BATF2
whose loss is mediated by unique metabolic cues in the TME and drives cancer immune escape.
STING–type-I interferon pathway regulates the immunogenicity of several cancer types, including head and neck squamous cell carcinoma. Here the authors describe that glutamine metabolism in the tumour microenvironment dampens the STING–type-I interferon pathway by epigenetically silencing the expression of BATF2, which functions as a tumour suppressor.
Journal Article
From Subjective to Principled Single-Cell Data Analysis: Evaluating Annotation Behavior, Optimizing Integration, and Benchmarking Visualization Pipelines
2026
Over recent years, single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics have transformed transcriptomic research by enabling high-resolution gene expression profiling and making it possible to generate atlas-scale datasets efficiently and cost-effectively. Yet rigorous scRNA-seq analysis remains challenging because key tasks, including cell-type annotation, data integration, and visualization, are often affected by substantial methodological and practical limitations. Cell-type annotation relies on a flexible multi-step pipeline whose functions and parameter choices are often shaped by analysts’ judgment and expertise, introducing subjectivity that can reduce reproducibility. Integration across batches is essential for large-scale and multi-condition studies, but it requires balancing the removal of between-batch variation with the preservation of true cell identity. Existing integration workflows still lack principled annotation-free strategies for feature selection and hyperparameter tuning. Visualization of atlas-scale data is further complicated by computational demands and by the strong influence of preprocessing steps such as normalization and integration, which remain insufficiently evaluated in existing benchmark studies. This dissertation addresses these challenges through three complementary projects on subjectivity in annotation, principled integration optimization, and systematic benchmarking of large-scale visualization pipelines. My first project examines how graduate students enrolled in an advanced bioinformatics course at UCLA performed cell-type annotation. We also collected information on their pipeline-tuning practices, annotation results, and educational and research backgrounds. Our study revealed that participants performed well in identifying major cell types but often struggled with closely related cell subtypes. Participants who achieved higher cell-type annotation accuracy tended to incorporate data-quality control into analysis pipelines or have prior publications with single-cell analysis. Subjective choices of parameter values influence clustering results, and tuning beyond default settings typically improves clustering accuracy. However, because cell-type label assignment is driven largely by prespecified marker genes and subjective judgments of their expression, parameter tuning has limited impact on improving annotation accuracy. We also identified a confirmation bias: prior expectations about cell types influence subjective annotation decisions. For comparison, we evaluated an AI agent, Biomni, for automated cell-type annotation and found its performance worse than 70–90% of participants. Our findings underscore the importance of transparent reporting of analysis pipelines, including parameter choice justifications and prior expectations about cell types. My second project presents IntegrateRigor, a data-driven, method-agnostic framework that performs batch-stable gene selection and optimizes integration hyperparameters across state-of-the-art methods, without relying on cell identity annotations. IntegrateRigor selects genes whose expression patterns are stable across batches using a gene-wise likelihood-based batch stability score, excluding batch-sensitive genes that can bias batch alignment during integration. IntegrateRigor then identifies optimal integration results across methods and hyperparameters by defining a dataset-level integration score that balances between-batch variation removal against cell identity preservation. In a colorectal cancer single-cell and spatial transcriptomics dataset, IntegrateRigor revealed previously uncharacterized cancer-immune interface states in the tumor microenvironment that were masked by both underintegration under default settings and over-integration in previous literature. Across diverse single-cell and spatial transcriptomics datasets, IntegrateRigor consistently improves integration performance by balancing over-integration and under-integration. By transforming integration from a heuristic step into a statistically principled, dataset-adaptive procedure, IntegrateRigor improves the reproducibility and discovery power of large-scale single-cell analyses. My third project proposes a comprehensive benchmark study evaluating 180 method pipelines that combine state-of-the-art normalization, integration, and visualization methods. We assessed performance on five semi-synthetic atlas-scale scRNA-seq datasets that mimic real data and include ground truths. Each pipeline was evaluated using ten quantitative metrics that assess visualization accuracy and scalability. This benchmark provided systematic, evidence-based guidance for selecting appropriate methods for visualizing atlas-scale single-cell data.
Dissertation
IntegrateRigor: annotation-free integration optimization for cell identity recovery reveals cancer-immune interface niches
2026
Integrating single-cell and spatial transcriptomics data across batches is essential for recovering comparable cell identities-including cell types, subtypes, and states-as a prerequisite for downstream analyses in multi-condition and large-scale studies. This task remains challenging because between-batch variation removal often conflicts with cell identity preservation, and current methods typically rely on generic highly variable gene selection and lack principled metrics for hyperparameter tuning when cell identity annotations are unavailable. Together, these limitations often lead to over-integration, which merges biologically distinct cell identities, or under-integration, which leaves cells separated by batch rather than identity. Here we introduce IntegrateRigor, a data-driven, annotation-free, method-agnostic framework that optimizes integration specifically for reliable cell identity recovery across batches. IntegrateRigor first selects genes whose expression patterns are stable across batches using a gene-wise likelihood-based batch stability score, excluding batch-sensitive genes that can bias cell identity alignment during integration. It then identifies the optimal integration configuration across methods and hyperparameters by defining a dataset-level integration score that explicitly balances between-batch variation removal against cell identity preservation, without requiring prior annotations. In a colorectal cancer single-cell and spatial transcriptomics dataset, IntegrateRigor revealed previously uncharacterized cancer-immune interface niches in the tumor microenvironment that were masked by under-integration under default settings and by over-integration in previous literature. Across diverse datasets spanning multiple sources of between-batch variation, IntegrateRigor consistently improved cell identity recovery by mitigating both over-integration and under-integration across five state-of-the-art integration methods. By transforming integration from a heuristic preprocessing step into a statistically principled, dataset-adaptive procedure for cell identity recovery, IntegrateRigor improves the reproducibility and biological discovery power of large-scale single-cell and spatial transcriptomics analyses.
Journal Article
CIPHER: An end-to-end framework for designing optimized aggregated spatial transcriptomics experiments
2026
Most imaging-based spatial transcriptomics methods measure individual genes, which limits scalability and typically requires integration with scRNA-seq to recover full cellular states. Recent approaches such as CISI, FISHnCHIPs, and ATLAS address this limitation by measuring aggregate transcriptional signatures, where multiple genes are pooled into each channel to increase throughput. While aggregate measurements improve scalability, they shift the problem from gene selection to feature design. For effective integration with scRNA-seq, these signatures must be not only discriminative in transcriptional space but also straightforward to measure, with balanced signal, sufficient dynamic range, and robustness to experimental noise. By optimizing decoding accuracy in isolation, existing methods leave substantial performance on the table.
We present CIPHER (Cell Identity Projection using Hybridization Encoding Rules), a neural-network framework that jointly optimizes the experimental encoding matrix, i.e., the way that genes are aggregated to signatures, and the downstream cell embedding. CIPHER integrates the physical limits of imaging assays directly into its loss function, shaping the latent space to maximize discriminability while maintaining robustness to measurement noise and signal constraints. Using a large-scale mouse brain scRNA-seq reference, we show that CIPHER-designed encodings yield latent spaces with improved cell-type separability, uniform signal utilization, and greater resilience to hybridization variability, resulting in higher decoding accuracy from both simulated and experimental data.
CIPHER formulates aggregate signature design as a joint optimization problem over decoding accuracy and experimental measurability. This enables systematic, scRNA-seq-aligned feature design for scalable spatial transcriptomics based on aggregate measurements.
Code and documentation are available at https://github.com/wollmanlab/Design/ .
Spatial transcriptomics reveals how cells are organized within tissues by mapping where genes are expressed. To achieve both scale and resolution, many approaches now combine spatial imaging with single-cell RNA-seq references to reconstruct complete transcriptomes
. New methods such as CISI, FISHnCHIPs, and ATLAS accelerate this process by measuring combinations of co-expressed genes rather than each gene individually. These aggregate measurements simplify experiments but introduce a new challenge: deciding which genes to combine so that the resulting features are both experimentally reliable and computationally informative for integration with scRNA-seq data. We developed CIPHER, a computational framework that learns how to design and decode these aggregate measurements optimally. By integrating experimental constraints and decoding accuracy into a unified neural-network model, CIPHER provides a principled approach for designing signature-based spatial transcriptomics experiments that enable efficient and accurate reconstruction of cellular transcriptomes.
Journal Article
Categorization and analysis of 14 computational methods for estimating cell potency from single-cell RNA-seq data
2024
In single-cell RNA sequencing (scRNA-seq) analysis, a key challenge is inferring hidden cellular dynamics from static cell snapshots. Various computational methods have been developed to address this, focusing on perspectives like pseudotime trajectories, RNA velocities, and estimating the differentiation potential of cells, often referred to as \"cell potency.\" This review summarizes 14 methods for defining cell potency from scRNA-seq data, categorizing them into average-based, entropy-based, and correlation-based methods based on how they summarize gene expression levels into a potency measure. We highlight the key similarities and differences within and between these categories, offering a high-level intuition for each method. Additionally, we use unified mathematical notations to detail each method's methodology and summarize their usage complexities, including parameters, required inputs, and differences between published descriptions and software implementations. We conclude that cell potency estimation remains an open question without a consensus on the optimal approach, emphasizing the need for benchmark datasets and studies. This review aims to provide a foundation for future benchmark studies, while also addressing the broader challenge of comparing methods that infer cellular dynamics from scRNA-seq data through various perspectives, including pseudotime trajectories, RNA velocities, and cell potency.
Supervised Capacity Preserving Mapping: A Clustering Guided Visualization Method for scRNAseq data
2021
The rapid development of scRNA-seq technologies enables us to explore the transcriptome at the cell level in a large scale. Recently, various computational methods have been developed to analyze the scRNAseq data such as clustering and visualization. However, current visualization methods including t-SNE and UMAP are challenged by the limited accuracy of rendering the geometic relationship of populations with distinct functional states. Most visualization methods are unsupervised, leaving out information from the clustering results or given labels. This leads to the inaccurate depiction of the distances between the bona fide functional states. In particular, UMAP and t-SNE are not optimal to preserve the global geometric structure. They may result in a contradiction that clusters with near distance in the embedded dimensions are in fact further away in the original dimensions. Besides, UMAP and t-SNE cannot track the variance of clusters. Through the embedding of t-SNE and UMAP, the variance of a cluster is not only associated with the true variance but also is proportional to the sample size. We present supCPM, a robust supervised visualization method, which separates different clusters, preserves global structure, and tracks the cluster variance. Compared with six visualization methods using synthetic and real datasets, supCPM shows improved performance than other methods in preserving the global geometric structure and data variance. Overall, supCPM provides an enhanced visualization pipeline to assist the interpretation of functional transition and accurately depict population segregation. Competing Interest Statement The authors have declared no competing interest. Footnotes * https://github.com/zhiqianZ/supCPM.git
Novel Corneal Protein Biomarker Candidates Reveal Iron Metabolic Disturbance in High Myopia Eyes
by
Alzogool, Mohammad
,
Zhai, Chuannan
,
Deng, Baocheng
in
Biomarkers
,
Cell and Developmental Biology
,
Chromatography
2021
Myopia is a major public health concern with increasing global prevalence and is the leading cause of vision loss and complications. The potential role of the cornea, a substantial component of refractive power and the protective fortress of the eye, has been underestimated in the development of myopia. Our study acquired corneal stroma tissues from myopic patients undergoing femtosecond laser-assisted small incision lenticule extraction (SMILE) surgery and investigated the differential expression of circulating proteins between subjects with low and high myopia by means of high-throughput proteomic approaches—the quantitative tandem mass tag (TMT) labeling method and parallel reaction monitoring (PRM) validation. Across all corneal stroma tissue samples, a total of 2,455 proteins were identified qualitatively and quantitatively, 103 of which were differentially expressed between those with low and high myopia. The differentially abundant proteins (DAPs) between the groups of stroma samples mostly demonstrated catalytic activity and molecular function regulator and transporter activity and participated in metabolic processes, biological regulation, response to stimulus, and so forth. Pathway enrichment showed that mineral absorption, ferroptosis, and HIF-1 signaling pathways were activated in the human myopic cornea. Furthermore, TMT analysis and PRM validation revealed that the expression of ferritin light chain (FTL, P02792) and ferritin heavy chain (FTH1, P02794) was negatively associated with myopia development, while the expression of serotransferrin (TF, P02787) was positively related to myopia status. Overall, our results indicated that subjects with low and high myopia could have different proteomic profiles or signatures in the cornea. These findings revealed disturbances in iron metabolism and corneal oxidative stress in the more myopic eyes. Iron metabolic proteins could serve as an essential modulator in the pathogenesis of myopia.
Journal Article
A stable rechargeable aqueous Zn-Air battery enabled by heterogeneous MoS2 cathode catalysts
2022
Aqueous rechargeable zinc (Zn)–air batteries have recently attracted extensive research interest due to their low cost, environmental benignity, safety, and high energy density. However, the sluggish kinetics of oxygen (O2) evolution reaction (OER) and the oxygen reduction reaction (ORR) of cathode catalysts in the batteries result in the high over-potential that impedes the practical application of Zn–air batteries. Here, we report a stable rechargeable aqueous Zn–air battery by use of a heterogeneous two-dimensional molybdenum sulfide (2D MoS2) cathode catalyst that consists of a heterogeneous interface and defects-embedded active edge sites. Compared to commercial Pt/C-RuO2, the low cost MoS2 cathode catalyst shows decent oxygen evolution and acceptable oxygen reduction catalytic activity. The assembled aqueous Zn–air battery using hybrid MoS2 catalysts demonstrates a specific capacity of 330 mAh g−1 and a durability of 500 cycles (~180 h) at 0.5 mA cm−2. In particular, the hybrid MoS2 catalysts outperform commercial Pt/C in the practically meaningful high-current region (>5 mA cm−2). This work paves the way for research on improving the performance of aqueous Zn–air batteries by constructing their own heterogeneous surfaces or interfaces instead of constructing bifunctional catalysts by compounding other materials.
Journal Article
The susceptibility of SERPINE1 rs1799889 SNP in diabetic vascular complications: a meta-analysis of fifty-one case-control studies
2021
Background
The serine protease inhibitor-1 (SERPINE1) rs1799889 single nucleotide polymorphism (SNP) has been constantly associated with diabetes mellitus (DM) and its vascular complications. The aim of this meta-analysis was to evaluate this association with combined evidences.
Methods
The systematic search was performed for studies published up to March 2021 which assess the associations between SERPINE1 rs1799889 SNP and the risks of DM, diabetic retinopathy (DR), diabetic cardiovascular disease (CVD) and diabetic nephropathy (DN). Only case-control studies were identified, and the linkage between SERPINE1 rs1799889 polymorphism and diabetic vascular risks were evaluated using genetic models.
Results
51 comparisons were enrolled. The results revealed a significant association with diabetes risk in overall population (allelic: OR = 1.34, 95 % CI = 1.14–1.57, homozygous: OR = 1.66, 95 % CI = 1.23–2.14, heterozygous: OR = 1.35, 95 % CI = 1.08–1.69, dominant: OR = 1.49, 95 % CI = 1.18–1.88, recessive: OR = 1.30, 95 % CI = 1.06–1.59) as well as in Asian descents (allelic: OR = 1.45, 95 % CI = 1.16–1.82, homozygous: OR = 1.88, 95 % CI = 1.29–2.75, heterozygous: OR = 1.47, 95 % CI = 1.08-2.00, dominant: OR = 1.64, 95 % CI = 1.21–2.24, recessive: OR = 1.46, 95 % CI = 1.09–1.96). A significant association was observed with DR risk (homozygous: OR = 1.25, 95 % CI = 1.01–1.56, recessive: OR = 1.20, 95 % CI = 1.01–1.43) for overall population, as for the European subgroup (homozygous: OR = 1.32, 95 % CI = 1.02–1.72, recessive: OR = 1.38, 95 % CI = 1.11–1.71). A significant association were shown with DN risk for overall population (allelic: OR = 1.48, 95 % CI = 1.15–1.90, homozygous: OR = 1.92, 95 % CI = 1.26–2.95, dominant: OR = 1.41, 95 % CI = 1.01–1.97, recessive: OR = 1.78, 95 % CI = 1.27–2.51) and for Asian subgroup (allelic: OR = 1.70, 95 % CI = 1.17–2.47, homozygous: OR = 2.46, 95 % CI = 1.30–4.66, recessive: OR = 2.24, 95 % CI = 1.40–3.59) after ethnicity stratification. No obvious association was implied with overall diabetic CVD risk in any genetic models, or after ethnicity stratification.
Conclusions
SERPINE1 rs1799889 4G polymorphism may outstand for serving as a genetic synergistic factor in overall DM and DN populations, positively for individuals with Asian descent. The association of SERPINE1 rs1799889 SNP and DR or diabetic CVD risks was not revealed.
Journal Article