Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
119
result(s) for
"631/208/457/649/2157"
Sort by:
Variant calling and benchmarking in an era of complete human genome sequences
by
Salit, Marc
,
Miga, Karen H
,
Zook, Justin M
in
Deep learning
,
DNA sequencing
,
Genetic diversity
2023
Genetic variant calling from DNA sequencing has enabled understanding of germline variation in hundreds of thousands of humans. Sequencing technologies and variant-calling methods have advanced rapidly, routinely providing reliable variant calls in most of the human genome. We describe how advances in long reads, deep learning, de novo assembly and pangenomes have expanded access to variant calls in increasingly challenging, repetitive genomic regions, including medically relevant regions, and how new benchmark sets and benchmarking methods illuminate their strengths and limitations. Finally, we explore the possible future of more complete characterization of human genome variation in light of the recent completion of a telomere-to-telomere human genome reference assembly and human pangenomes, and we consider the innovations needed to benchmark their newly accessible repetitive regions and complex variants.Variant calling is the process of identifying genetic variants, which is important for characterizing population genetic diversity and for identifying disease-associated variants in clinical sequencing projects. In this Review, the authors discuss the state-of-the-art in variant calling, focusing on challenging types of genetic variants, advances in both sequencing technologies and computational pipelines, and benchmarking strategies to assess the robustness of variant-calling strategies.
Journal Article
Mapping and characterization of structural variation in 17,795 human genomes
2020
A key goal of whole-genome sequencing for studies of human genetics is to interrogate all forms of variation, including single-nucleotide variants, small insertion or deletion (indel) variants and structural variants. However, tools and resources for the study of structural variants have lagged behind those for smaller variants. Here we used a scalable pipeline
1
to map and characterize structural variants in 17,795 deeply sequenced human genomes. We publicly release site-frequency data to create the largest, to our knowledge, whole-genome-sequencing-based structural variant resource so far. On average, individuals carry 2.9 rare structural variants that alter coding regions; these variants affect the dosage or structure of 4.2 genes and account for 4.0–11.2% of rare high-impact coding alleles. Using a computational model, we estimate that structural variants account for 17.2% of rare alleles genome-wide, with predicted deleterious effects that are equivalent to loss-of-function coding alleles; approximately 90% of such structural variants are noncoding deletions (mean 19.1 per genome). We report 158,991 ultra-rare structural variants and show that 2% of individuals carry ultra-rare megabase-scale structural variants, nearly half of which are balanced or complex rearrangements. Finally, we infer the dosage sensitivity of genes and noncoding elements, and reveal trends that relate to element class and conservation. This work will help to guide the analysis and interpretation of structural variants in the era of whole-genome sequencing.
Structural variants in more than 17,000 human genomes are mapped and characterized using whole-genome sequencing, showing how this type of variation contributes to rare deleterious coding and noncoding alleles.
Journal Article
Schizophrenia risk from complex variation of complement component 4
by
Davis, Avery
,
Van Doren, Vanessa
,
Kamitaki, Nolan
in
631/208
,
631/208/457/649/2157
,
631/208/728
2016
Schizophrenia is a heritable brain illness with unknown pathogenic mechanisms. Schizophrenia’s strongest genetic association at a population level involves variation in the major histocompatibility complex (MHC) locus, but the genes and molecular mechanisms accounting for this have been challenging to identify. Here we show that this association arises in part from many structurally diverse alleles of the complement component 4 (
C4
) genes. We found that these alleles generated widely varying levels of
C4A
and
C4B
expression in the brain, with each common
C4
allele associating with schizophrenia in proportion to its tendency to generate greater expression of
C4A
. Human C4 protein localized to neuronal synapses, dendrites, axons, and cell bodies. In mice, C4 mediated synapse elimination during postnatal development. These results implicate excessive complement activity in the development of schizophrenia and may help explain the reduced numbers of synapses in the brains of individuals with schizophrenia.
WebSchizophrenia is associated with genetic variation at the major histocompatibility complex locus; this study reveals that alleles at this locus associate with schizophrenia in proportion to their tendency to generate greater expression of complement component 4 (
C4A
) genes and that C4 promotes the elimination of synpases.
The genetics of schizophrenia
The strongest genetic association found in schizophrenia is its association to genetic markers across the major histocompatibility complex (MHC) locus, first described in three
Nature
papers in 2009. The association signal at the MHC is extremely complex. Here Steven McCarroll and colleagues report a dissection of the MHC association to schizophrenia. They find a strong contribution from many structurally diverse alleles of the complement component 4 (
C4
) genes. The linkage was higher for
C4
alleles that promoted greater expression of
C4A
, measured in the brain tissues of adult post-mortem donors with or without schizophrenia. The authors suggest that C4 may work with other components of the classical complement cascade to promote synaptic pruning, and demonstrate that C4 mediates synaptic refinement in a mouse model.
Journal Article
GraphTyper2 enables population-scale genotyping of structural variation using pangenome graphs
by
Gudbjartsson, Daniel F.
,
Kristmundsdottir, Snaedis
,
Stefansson, Kari
in
45/23
,
631/114/794
,
631/208/457/649/2157
2019
Analysis of sequence diversity in the human genome is fundamental for genetic studies. Structural variants (SVs) are frequently omitted in sequence analysis studies, although each has a relatively large impact on the genome. Here, we present GraphTyper2, which uses pangenome graphs to genotype SVs and small variants using short-reads. Comparison to the syndip benchmark dataset shows that our SV genotyping is sensitive and variant segregation in families demonstrates the accuracy of our approach. We demonstrate that incorporating public assembly data into our pipeline greatly improves sensitivity, particularly for large insertions. We validate 6,812 SVs on average per genome using long-read data of 41 Icelanders. We show that GraphTyper2 can simultaneously genotype tens of thousands of whole-genomes by characterizing 60 million small variants and half a million SVs in 49,962 Icelanders, including 80 thousand SVs with high-confidence.
Structural variants may be omitted in sequence analysis despite their importance in genome variation and phenotypic impact. Here the authors present GraphTyper2, which uses pangenome graphs to genotype structural variants using short-reads and can be applied in large-scale sequencing studies.
Journal Article
ERα-associated translocations underlie oncogene amplifications in breast cancer
by
Viswanadham, Vinayak V.
,
Pellman, David
,
Chu, Chong
in
45/23
,
631/208/457/649/2157
,
631/67/1347
2023
Focal copy-number amplification is an oncogenic event. Although recent studies have revealed the complex structure
1
–
3
and the evolutionary trajectories
4
of oncogene amplicons, their origin remains poorly understood. Here we show that focal amplifications in breast cancer frequently derive from a mechanism—which we term translocation–bridge amplification—involving inter-chromosomal translocations that lead to dicentric chromosome bridge formation and breakage. In 780 breast cancer genomes, we observe that focal amplifications are frequently connected to each other by inter-chromosomal translocations at their boundaries. Subsequent analysis indicates the following model: the oncogene neighbourhood is translocated in G1 creating a dicentric chromosome, the dicentric chromosome is replicated, and as dicentric sister chromosomes segregate during mitosis, a chromosome bridge is formed and then broken, with fragments often being circularized in extrachromosomal DNAs. This model explains the amplifications of key oncogenes, including
ERBB2
and
CCND1
. Recurrent amplification boundaries and rearrangement hotspots correlate with oestrogen receptor binding in breast cancer cells. Experimentally, oestrogen treatment induces DNA double-strand breaks in the oestrogen receptor target regions that are repaired by translocations, suggesting a role of oestrogen in generating the initial translocations. A pan-cancer analysis reveals tissue-specific biases in mechanisms initiating focal amplifications, with the breakage–fusion–bridge cycle prevalent in some and the translocation–bridge amplification in others, probably owing to the different timing of DNA break repair. Our results identify a common mode of oncogene amplification and propose oestrogen as its mechanistic origin in breast cancer.
An analysis of 780 breast cancer genomes shows that focal amplifications are frequently preceded by dicentric chromosome formation from inter-chromosomal translocations associated with oestrogen receptor binding, which leads to chromosome bridge formation and breakage, initiating the amplification process.
Journal Article
Refined genetic maps reveal sexual dimorphism in human meiotic recombination at multiple scales
by
Campbell, Christopher L.
,
Bhérer, Claude
,
Auton, Adam
in
631/181/457
,
631/208/457
,
631/208/457/649/2157
2017
In humans, males have lower recombination rates than females over the majority of the genome, but the opposite is usually true near the telomeres. These broad-scale differences have been known for decades, yet little is known about differences at the fine scale. By combining data sets, we have collected recombination events from over 100,000 meioses and have constructed sex-specific genetic maps at a previously unachievable resolution. Here we show that, although a substantial fraction of the genome shows some degree of sexually dimorphic recombination, the vast majority of hotspots are shared between the sexes, with only a small number of putative sex-specific hotspots. Wavelet analysis indicates that most of the differences can be attributed to the fine scale, and that variation in rate between the sexes can mostly be explained by differences in hotspot magnitude, rather than location. Nonetheless, known recombination-associated genomic features, such as THE1B repeat elements, show systematic differences between the sexes.
It is known that males have lower recombination rates than females over much of the genome but little is known about differences at a fine scale. Here the authors combine data from over 100,000 meioses and show that the majority of differences can be explained by variation in hotspot magnitude.
Journal Article
Jasmine and Iris: population-scale structural variant comparison and analysis
2023
The availability of long reads is revolutionizing studies of structural variants (SVs). However, because SVs vary across individuals and are discovered through imprecise read technologies and methods, they can be difficult to compare. Addressing this, we present Jasmine and Iris (
https://github.com/mkirsche/Jasmine/
), for fast and accurate SV refinement, comparison and population analysis. Using an SV proximity graph, Jasmine outperforms six widely used comparison methods, including reducing the rate of Mendelian discordance in trio datasets by more than fivefold, and reveals a set of high-confidence de novo SVs confirmed by multiple technologies. We also present a unified callset of 122,813 SVs and 82,379 indels from 31 samples of diverse ancestry sequenced with long reads. We genotype these variants in 1,317 samples from the 1000 Genomes Project and the Genotype-Tissue Expression project with DNA and RNA-sequencing data and assess their widespread impact on gene expression, including within medically relevant genes.
An optimized pipeline for improved inference and analysis of structural variants (SVs) has been developed, which uses Iris for refining breakpoints and sequences, and Jasmine for comparing SV calls at population scale.
Journal Article
Population genomics of bacterial host adaptation
by
Guttman, David S
,
J Ross Fitzgerald
,
Sheppard, Samuel K
in
Bacteria
,
Food security
,
Genomic analysis
2018
Some bacteria can transfer to new host species, and this poses a risk to human health. Indeed, an estimated 60% of all human pathogens have originated from other animal species. Similarly, human-to-animal transitions are recognized as a major threat to sustainable livestock production, and emerging pathogens impose an increasing burden on crop yield and global food security. Recent advances in high-throughput sequencing technologies have enabled comparative genomic analyses of bacterial populations from multiple hosts. Such studies are providing new insights into the evolutionary processes that underpin the establishment of bacteria in new host niches. A better understanding of the genetic and mechanistic basis for bacterial host adaptation may reveal novel targets for controlling infection or inform the design of approaches to limit the emergence of new pathogens.
Journal Article
CaSpER identifies and visualizes CNV events by integrative analysis of single-cell or bulk RNA-sequencing data
by
Harmanci, Arif O.
,
Serin Harmanci, Akdes
,
Zhou, Xiaobo
in
45/91
,
631/114
,
631/208/457/649/2157
2020
RNA sequencing experiments generate large amounts of information about expression levels of genes. Although they are mainly used for quantifying expression levels, they contain much more biologically important information such as copy number variants (CNVs). Here, we present CaSpER, a signal processing approach for identification, visualization, and integrative analysis of focal and large-scale CNV events in multiscale resolution using either bulk or single-cell RNA sequencing data. CaSpER integrates the multiscale smoothing of expression signal and allelic shift signals for CNV calling. The allelic shift signal measures the loss-of-heterozygosity (LOH) which is valuable for CNV identification. CaSpER employs an efficient methodology for the generation of a genome-wide B-allele frequency (BAF) signal profile from the reads and utilizes it for correction of CNVs calls. CaSpER increases the utility of RNA-sequencing datasets and complements other tools for complete characterization and visualization of the genomic and transcriptomic landscape of single cell and bulk RNA sequencing data.
RNA-sequencing is mostly used to assess gene expression; however, it can also give information about genetic variants. Here, the authors present CaSpER, a statistical framework that utilises RNA-sequencing reads to identify and visualise CNV events by integrating transcriptome-wide expression and allelic shift profiles.
Journal Article