Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
67
result(s) for
"Zhan, Jianan"
Sort by:
Transcriptome analysis reveals dysregulation of innate immune response genes and neuronal activity-dependent genes in autism
by
Gupta, Simone
,
Ellis, Shannon E.
,
West, Andrew B.
in
631/208/212/2019
,
692/420/2780/262
,
692/699/476/1373
2014
Recent studies of genomic variation associated with autism have suggested the existence of extreme heterogeneity. Large-scale transcriptomics should complement these results to identify core molecular pathways underlying autism. Here we report results from a large-scale RNA sequencing effort, utilizing region-matched autism and control brains to identify neuronal and microglial genes robustly dysregulated in autism cortical brain. Remarkably, we note that a gene expression module corresponding to M2-activation states in microglia is negatively correlated with a differentially expressed neuronal module, implicating dysregulated microglial responses in concert with altered neuronal activity-dependent genes in autism brains. These observations provide pathways and candidate genes that highlight the interplay between innate immunity and neuronal activity in the aetiology of autism.
Autism spectrum disorder (ASD) is a common, highly heritable neurodevelopmental condition characterized by marked genetic heterogeneity. In this study, the authors use RNA sequencing analyses to characterize differences in the transcriptome between autistic and typically developing brains.
Journal Article
An ensemble penalized regression method for multi-ancestry polygenic risk prediction
by
Zhang, Haoyu
,
Chatterjee, Nilanjan
,
Ma, Cheng
in
631/208/205
,
631/208/727/2000
,
Bayes Theorem
2024
Great efforts are being made to develop advanced polygenic risk scores (PRS) to improve the prediction of complex traits and diseases. However, most existing PRS are primarily trained on European ancestry populations, limiting their transferability to non-European populations. In this article, we propose a novel method for generating multi-ancestry Polygenic Risk scOres based on enSemble of PEnalized Regression models (PROSPER). PROSPER integrates genome-wide association studies (GWAS) summary statistics from diverse populations to develop ancestry-specific PRS with improved predictive power for minority populations. The method uses a combination of
L
1
(lasso) and
L
2
(ridge) penalty functions, a parsimonious specification of the penalty parameters across populations, and an ensemble step to combine PRS generated across different penalty parameters. We evaluate the performance of PROSPER and other existing methods on large-scale simulated and real datasets, including those from 23andMe Inc., the Global Lipids Genetics Consortium, and All of Us. Results show that PROSPER can substantially improve multi-ancestry polygenic prediction compared to alternative methods across a wide variety of genetic architectures. In real data analyses, for example, PROSPER increased out-of-sample prediction R
2
for continuous traits by an average of 70% compared to a state-of-the-art Bayesian method (PRS-CSx) in the African ancestry population. Further, PROSPER is computationally highly scalable for the analysis of large SNP contents and many diverse populations.
Great efforts are being made to develop advanced polygenic risk scores (PRS) to improve the prediction of complex traits and diseases. However most existing PRS are primarily trained on European ancestry populations, limiting their transferability to non-European populations. Here the authors propose a new multi-ancestry PRS method, PROSPER, to reduce disparity of PRS performance across ancestry groups.
Journal Article
A new method for multiancestry polygenic prediction improves performance across diverse populations
2023
Polygenic risk scores (PRSs) increasingly predict complex traits; however, suboptimal performance in non-European populations raise concerns about clinical applications and health inequities. We developed CT-SLEB, a powerful and scalable method to calculate PRSs, using ancestry-specific genome-wide association study summary statistics from multiancestry training samples, integrating clumping and thresholding, empirical Bayes and superlearning. We evaluated CT-SLEB and nine alternative methods with large-scale simulated genome-wide association studies (~19 million common variants) and datasets from 23andMe, Inc., the Global Lipids Genetics Consortium, All of Us and UK Biobank, involving 5.1 million individuals of diverse ancestry, with 1.18 million individuals from four non-European populations across 13 complex traits. Results demonstrated that CT-SLEB significantly improves PRS performance in non-European populations compared with simple alternatives, with comparable or superior performance to a recent, computationally intensive method. Moreover, our simulation studies offered insights into sample size requirements and SNP density effects on multiancestry risk prediction.
CT-SLEB, a powerful and scalable method, improves the performance of multiancestry polygenic prediction by generating polygenic risk scores based on GWAS summary statistics in diverse populations.
Journal Article
CYP2C19 Allele Frequencies in Over 2.2 Million Direct‐to‐Consumer Genetics Research Participants and the Potential Implication for Prescriptions in a Large Health System
2020
Understanding the prevalence of clinically relevant pharmacogenetic variants using large unselected populations is critical for gauging the potential clinical impact of widespread preemptive pharmacogenetic testing. To this end, we assessed the frequencies and ethnic distribution of the three most common CYP2C19 alleles (*2, *3, and *17) in 2.29 million direct‐to‐consumer genetics research participants (23andMe, Sunnyvale, CA). The overall frequencies of *2, *3, and *17 were 15.2%, 0.3%, and 20.4%, respectively, but varied by ethnicity. The most common variant diplotypes were *1/*17 at 26% and *1/*2 at 19.4%. The less common *2/*17, *17/*17, and *2/*2 genotypes occurred at 6.0%, 4.4%, and 2.5%, respectively. Overall, 58.3% of participants had at least one increased‐function or no‐function CYP2C19 allele. To better understand how this high frequency might impact a real patient population, we examined the prescription rates (Rx) of high‐pharmacogenetic‐risk medications metabolized by CYP2C19 using the University of California at San Francisco (UCSF) health system’s anonymized database of over 1.25 million patients. Between 2012 and 2019, a total of 151,068 UCSF patients (15.8%) representing 5 self‐reported ethnicities were prescribed one or more high‐pharmacogenetic‐risk CYP2C19 medications: proton pump inhibitors (145,243 Rx), three selective serotonin reuptake inhibitor antidepressants (54,463 Rx), clopidogrel (14,376 Rx), and voriconazole (2,303 Rx).
Journal Article
Priors, population sizes, and power in genome-wide hypothesis tests
2023
Background
Genome-wide tests, including genome-wide association studies (GWAS) of germ-line genetic variants, driver tests of cancer somatic mutations, and transcriptome-wide association tests of RNAseq data, carry a high multiple testing burden. This burden can be overcome by enrolling larger cohorts or alleviated by using prior biological knowledge to favor some hypotheses over others. Here we compare these two methods in terms of their abilities to boost the power of hypothesis testing.
Results
We provide a quantitative estimate for progress in cohort sizes and present a theoretical analysis of the power of oracular hard priors: priors that select a subset of hypotheses for testing, with an oracular guarantee that all true positives are within the tested subset. This theory demonstrates that for GWAS, strong priors that limit testing to 100–1000 genes provide less power than typical annual 20–40% increases in cohort sizes. Furthermore, non-oracular priors that exclude even a small fraction of true positives from the tested set can perform worse than not using a prior at all.
Conclusion
Our results provide a theoretical explanation for the continued dominance of simple, unbiased univariate hypothesis tests for GWAS: if a statistical question can be answered by larger cohort sizes, it should be answered by larger cohort sizes rather than by more complicated biased methods involving priors. We suggest that priors are better suited for non-statistical aspects of biology, such as pathway structure and causality, that are not yet easily captured by standard hypothesis tests.
Journal Article
1181 Prevalence of Celiac Disease-Compatible Human Leukocyte Antigen Haplotypes Across Ethnicities and Regions in the United States
2019
INTRODUCTION:The prevalence of celiac disease (CD) is widely variable throughout the United States (US), with a higher prevalence of disease in the Northeast. The reasons for this variability are unknown. In a prior study, we detected ethnic differences within the US, using a name-based algorithm, with prevalence in patients of Jewish ethnicity similar to the overall population, and lower in persons of East Asian ethnicity. CD etiology is dependent on human leukocyte antigen (HLA) haplotype. Typically, either HLA DQ2.5 or DQ8 is required (but not sufficient) for the development of CD, with DQ2.5 being the highest-risk haplotype. To date, no study has characterized regional or ethnic differences in the frequency of CD-compatible HLA haplotypes. Thus, we aimed to measure the frequencies of DQ2.5 and DQ8 across regions and ethnicities in the US.METHODS:We assessed the frequencies of HLA DQ2.5 (DQA1*05:DQB1*02) and DQ8 (DQA1*03:DQB1*03) in an unselected group of genotyped individuals who have used direct-to-consumer genetic testing between 2013 and 2017. Eligible participants were 23andMe customers who consented to participate in research. We assayed two SNPs to classify individuals as DQ2.5 homozygous, DQ2.5 heterozygous, DQ2.5/DQ8, DQ8 homozygous, DQ8 heterozygous, and 0 detected variants. We compared the frequency of each haplotype across four regions of the US. Additionally, we used genome-wide array data to cluster participants into 8 categories that correlate highly with self-reported race and ethnicity, and compared the frequencies of these haplotypes across these ethnic categories (Table 1).RESULTS:Of 1,290,668 individuals studied, at least one CD-compatible haplotype was present in 38.7% of individuals, and this frequency was similar across the four US regions. The frequencies of DQ2.5 homozygotes were also similar across the Northeast, Midwest, South, and West (1.25%, 1.43%, 1.38%, and 1.37%, respectively). In contrast, frequencies differ across ethnic groups: the highest DQ2.5 and DQ8 frequencies were observed in European (12.01%) and Ashkenazi Jewish (16.39%) participants, respectively.CONCLUSION:Previously reported regional variability in CD prevalence in the US may not be due to differences in HLA-based susceptibility; rather, other genetic or environmental factors likely play a role in disease pathogenesis. In addition, these differences carry great significance in view of the development of HLA haplotype-specific non-dietary therapies for CD.Table 1.Distribution of DQ2.5 and DQ8 haplotype frequencies observed across ethnic categories
Journal Article
Explore the Relationship Between Genetic Variations and Phenotypes with Bayesian Approaches
2018
Genome-wide association studies (GWAS) have had great success in identifying human genetic variants associated with human traits. With recent developments in high throughput biology, immense amount of data have been generated, thus calling for novel statistical and computational approaches to be developed and draw biological meaningful conclusions. A current direction for GWAS method development has been to use Bayesian approaches, where prior beliefs of variant effects are incorporated into test statistics, to boost the power to detect real associations. With previous success in developing Bayesian-based GWAS method for single phenotype, in this work the Bayesian idea is extended to multiple phenotypes, aiming at developing a method that detects pleiotropic genome-wide associations. Alongside with the method development, analytical simulations were also performed to investigate into the possible power gain by using such Bayesian approaches, as well as to understand how different factors influence the behavior of Bayesian-based GWAS methods. Many variants are pleiotropic, and discovery of these variants could help reveal disease mechanisms, suggest new therapeutic options. Therefore, we developed a pleiotropic GWAS method based on Bayesian framework, SNP And Pleiotropic PHenotype Organization (SAPPHO), which learns pleiotropy using identified associations to discover additional associations with shared patterns. SAPPHO was applied on two sets of real data: 1. Atherosclerosis Risk in Communities (ARIC) study of 8,000 individuals, whose gold-standard associations were provided by meta-analysis of 40,000 to 100,000 individuals from the Cohorts for Heart and Aging Research in Genomic Epidemiology (CHARGE) consortium; 2. Cancer phenotypes from UK Biobank project, consisting several hundred to 15,000 individuals, with gold-standard obtained from GWAS catalog. For both data sets, SAPPHO was able to detect additional associations that were not detected with the conventional univariate test, and boost power when different variants follow the same association patterns. Bayesian approaches boost power for GWAS through alleviating burdens from multiple hypothesis testings, which is usually on the scale of thousands to millions. Intuitively, by making use of prior probabilities that bias favored sets thought to be enriched for significant findings, power for detecting true associations could be increased. Therefore, an analytical study was conducted here to see theoretically to what extent power could gain by using such approaches, and how does this gain depend on different factors. By calculating test power assuming perfect knowledge of a prior distribution, the population size increase required to provided the same boost without a prior was obtained, and it is shown that population size is exponentially more important than prior, providing a rigorous proof for the lack of use for prior-based GWAS methods.
Dissertation
A new method for multi-ancestry polygenic prediction improves performance across diverse populations
2023
Polygenic risk scores (PRS) increasingly predict complex traits, however, suboptimal performance in non-European populations raise concerns about clinical applications and health inequities. We developed CT-SLEB, a powerful and scalable method to calculate PRS using ancestry-specific GWAS summary statistics from multi-ancestry training samples, integrating clumping and thresholding, empirical Bayes and super learning. We evaluate CT-SLEB and nine-alternatives methods with large-scale simulated GWAS (∼19 million common variants) and datasets from 23andMe Inc., the Global Lipids Genetics Consortium, All of Us and UK Biobank involving 5.1 million individuals of diverse ancestry, with 1.18 million individuals from four non-European populations across thirteen complex traits. Results demonstrate that CT-SLEB significantly improves PRS performance in non-European populations compared to simple alternatives, with comparable or superior performance to a recent, computationally intensive method. Moreover, our simulation studies offer insights into sample size requirements and SNP density effects on multi-ancestry risk prediction.
An Ensemble Penalized Regression Method for Multi-ancestry Polygenic Risk Prediction
2024
Great efforts are being made to develop advanced polygenic risk scores (PRS) to improve the prediction of complex traits and diseases. However, most existing PRS are primarily trained on European ancestry populations, limiting their transferability to non-European populations. In this article, we propose a novel method for generating multi-ancestry Polygenic Risk scOres based on enSemble of PEnalized Regression models (PROSPER). PROSPER integrates genome-wide association studies (GWAS) summary statistics from diverse populations to develop ancestry-specific PRS with improved predictive power for minority populations. The method uses a combination of ℒ1 (lasso) and ℒ2 (ridge) penalty functions, a parsimonious specification of the penalty parameters across populations, and an ensemble step to combine PRS generated across different penalty parameters. We evaluate the performance of PROSPER and other existing methods on large-scale simulated and real datasets, including those from 23andMe Inc., the Global Lipids Genetics Consortium, and All of Us. Results show that PROSPER can substantially improve multi-ancestry polygenic prediction compared to alternative methods across a wide variety of genetic architectures. In real data analyses, for example, PROSPER increased out-of-sample prediction R2 for continuous traits by an average of 70% compared to a state-of-the-art Bayesian method (PRS-CSx) in the African ancestry population. Further, PROSPER is computationally highly scalable for the analysis of large SNP contents and many diverse populations.Great efforts are being made to develop advanced polygenic risk scores (PRS) to improve the prediction of complex traits and diseases. However, most existing PRS are primarily trained on European ancestry populations, limiting their transferability to non-European populations. In this article, we propose a novel method for generating multi-ancestry Polygenic Risk scOres based on enSemble of PEnalized Regression models (PROSPER). PROSPER integrates genome-wide association studies (GWAS) summary statistics from diverse populations to develop ancestry-specific PRS with improved predictive power for minority populations. The method uses a combination of ℒ1 (lasso) and ℒ2 (ridge) penalty functions, a parsimonious specification of the penalty parameters across populations, and an ensemble step to combine PRS generated across different penalty parameters. We evaluate the performance of PROSPER and other existing methods on large-scale simulated and real datasets, including those from 23andMe Inc., the Global Lipids Genetics Consortium, and All of Us. Results show that PROSPER can substantially improve multi-ancestry polygenic prediction compared to alternative methods across a wide variety of genetic architectures. In real data analyses, for example, PROSPER increased out-of-sample prediction R2 for continuous traits by an average of 70% compared to a state-of-the-art Bayesian method (PRS-CSx) in the African ancestry population. Further, PROSPER is computationally highly scalable for the analysis of large SNP contents and many diverse populations.
Journal Article
MUSSEL: Enhanced Bayesian Polygenic Risk Prediction Leveraging Information across Multiple Ancestry Groups
2023
Polygenic risk scores (PRS) are now showing promising predictive performance on a wide variety of complex traits and diseases, but there exists a substantial performance gap across different populations. We propose MUSSEL, a method for ancestry-specific polygenic prediction that borrows information in the summary statistics from genome-wide association studies (GWAS) across multiple ancestry groups. MUSSEL conducts Bayesian hierarchical modeling under a MUltivariate Spike-and-Slab model for effect-size distribution and incorporates an Ensemble Learning step using super learner to combine information across different tuning parameter settings and ancestry groups. In our simulation studies and data analyses of 16 traits across four distinct studies, totaling 5.7 million participants with a substantial ancestral diversity, MUSSEL shows promising performance compared to alternatives. The method, for example, has an average gain in prediction R2 across 11 continuous traits of 40.2% and 49.3% compared to PRS-CSx and CT-SLEB, respectively, in the African Ancestry population. The best-performing method, however, varies by GWAS sample size, target ancestry, underlying trait architecture, and the choice of reference samples for LD estimation, and thus ultimately, a combination of methods may be needed to generate the most robust PRS across diverse populations.Polygenic risk scores (PRS) are now showing promising predictive performance on a wide variety of complex traits and diseases, but there exists a substantial performance gap across different populations. We propose MUSSEL, a method for ancestry-specific polygenic prediction that borrows information in the summary statistics from genome-wide association studies (GWAS) across multiple ancestry groups. MUSSEL conducts Bayesian hierarchical modeling under a MUltivariate Spike-and-Slab model for effect-size distribution and incorporates an Ensemble Learning step using super learner to combine information across different tuning parameter settings and ancestry groups. In our simulation studies and data analyses of 16 traits across four distinct studies, totaling 5.7 million participants with a substantial ancestral diversity, MUSSEL shows promising performance compared to alternatives. The method, for example, has an average gain in prediction R2 across 11 continuous traits of 40.2% and 49.3% compared to PRS-CSx and CT-SLEB, respectively, in the African Ancestry population. The best-performing method, however, varies by GWAS sample size, target ancestry, underlying trait architecture, and the choice of reference samples for LD estimation, and thus ultimately, a combination of methods may be needed to generate the most robust PRS across diverse populations.
Journal Article