Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
76
result(s) for
"Lunter, Gerton"
Sort by:
A unified haplotype-based method for accurate and comprehensive variant calling
by
Wedge, David C.
,
Lunter, Gerton
,
Cooke, Daniel P.
in
631/114/2785/2302
,
631/114/794
,
631/553/2714
2021
Almost all haplotype-based variant callers were designed specifically for detecting common germline variation in diploid populations, and give suboptimal results in other scenarios. Here we present Octopus, a variant caller that uses a polymorphic Bayesian genotyping model capable of modeling sequencing data from a range of experimental designs within a unified haplotype-aware framework. Octopus combines sequencing reads and prior information to phase-called genotypes of arbitrary ploidy, including those with somatic mutations. We show that Octopus accurately calls germline variants in individuals, including single nucleotide variants, indels and small complex replacements such as microinversions. Using a synthetic tumor data set derived from clean sequencing data from a sample with known germline haplotypes and observed mutations in a large cohort of tumor samples, we show that Octopus is more sensitive to low-frequency somatic variation, yet calls considerably fewer false positives than other methods. Octopus also outputs realigned evidence BAM files to aid validation and interpretation.
Octopus detects germline and somatic variants with high sensitivity and accuracy.
Journal Article
DeepC: predicting 3D genome folding using megabase-scale transfer learning
by
Oudelaar, A. Marieke
,
Teh, Yee Whye
,
Lunter, Gerton
in
631/114/1305
,
631/114/2785
,
631/208/177
2020
Predicting the impact of noncoding genetic variation requires interpreting it in the context of three-dimensional genome architecture. We have developed deepC, a transfer-learning-based deep neural network that accurately predicts genome folding from megabase-scale DNA sequence. DeepC predicts domain boundaries at high resolution, learns the sequence determinants of genome folding and predicts the impact of both large-scale structural and single base-pair variations.
DeepC uses transfer learning-based deep neural networks for predicting genome folding from megabase-scale DNA sequence.
Journal Article
Insertions and Deletions: Computational Methods, Evolutionary Dynamics, and Biological Applications
2024
Abstract
Insertions and deletions constitute the second most important source of natural genomic variation. Insertions and deletions make up to 25% of genomic variants in humans and are involved in complex evolutionary processes including genomic rearrangements, adaptation, and speciation. Recent advances in long-read sequencing technologies allow detailed inference of insertions and deletion variation in species and populations. Yet, despite their importance, evolutionary studies have traditionally ignored or mishandled insertions and deletions due to a lack of comprehensive methodologies and statistical models of insertions and deletion dynamics. Here, we discuss methods for describing insertions and deletion variation and modeling insertions and deletions over evolutionary time. We provide practical advice for tackling insertions and deletions in genomic sequences and illustrate our discussion with examples of insertions and deletion-induced effects in human and other natural populations and their contribution to evolutionary processes. We outline promising directions for future developments in statistical methodologies that would allow researchers to analyze insertions and deletion variation and their effects in large genomic data sets and to incorporate insertions and deletions in evolutionary inference.
Journal Article
Sequencing of human genomes with nanopore technology
2019
Whole-genome sequencing (WGS) is becoming widely used in clinical medicine in diagnostic contexts and to inform treatment choice. Here we evaluate the potential of the Oxford Nanopore Technologies (ONT) MinION long-read sequencer for routine WGS by sequencing the reference sample NA12878 and the genome of an individual with ataxia-pancytopenia syndrome and severe immune dysregulation. We develop and apply a novel reference panel-free analytical method to infer and then exploit phase information which improves single-nucleotide variant (SNV) calling performance from otherwise modest levels. In the clinical sample, we identify and directly phase two non-synonymous de novo variants in
SAMD9L
, (OMIM #159550) inferring that they lie on the same paternal haplotype. Whilst consensus SNV-calling error rates from ONT data remain substantially higher than those from short-read methods, we demonstrate the substantial benefits of analytical innovation. Ongoing improvements to base-calling and SNV-calling methodology must continue for nanopore sequencing to establish itself as a primary method for clinical WGS.
Nanopore sequencing technology generates longer reads than current technologies, but with more errors. Here, the authors develop new analytical tools to improve accuracy and evaluate the potential of nanopore sequencing for clinical human genomics.
Journal Article
Integrating mapping-, assembly- and haplotype-based approaches for calling variants in clinical sequencing applications
2014
Gerton Lunter and colleagues report Platypus software, which combines a haplotype-based multi-sample variant caller with local sequence assembly in a Bayesian statistical framework. They demonstrate applications to exome and whole-genome data sets, to the identification
de novo
mutations in parent-offspring trios and to the genotyping of HLA loci.
High-throughput DNA sequencing technology has transformed genetic research and is starting to make an impact on clinical practice. However, analyzing high-throughput sequencing data remains challenging, particularly in clinical settings where accuracy and turnaround times are critical. We present a new approach to this problem, implemented in a software package called Platypus. Platypus achieves high sensitivity and specificity for SNPs, indels and complex polymorphisms by using local
de novo
assembly to generate candidate variants, followed by local realignment and probabilistic haplotype estimation. It is an order of magnitude faster than existing tools and generates calls from raw aligned read data without preprocessing. We demonstrate the performance of Platypus in clinically relevant experimental designs by comparing with SAMtools and GATK on whole-genome and exome-capture data, by identifying
de novo
variation in 15 parent-offspring trios with high sensitivity and specificity, and by estimating human leukocyte antigen genotypes directly from variant calls.
Journal Article
The Diversity and Molecular Evolution of B-Cell Receptors during Infection
2016
B-cell receptors (BCRs) are membrane-bound immunoglobulins that recognize and bind foreign proteins (antigens). BCRs are formed through random somatic changes of germline DNA, creating a vast repertoire of unique sequences that enable individuals to recognize a diverse range of antigens. After encountering antigen for the first time, BCRs undergo a process of affinity maturation, whereby cycles of rapid somatic mutation and selection lead to improved antigen binding. This constitutes an accelerated evolutionary process that takes place over days or weeks. Next-generation sequencing of the gene regions that determine BCR binding has begun to reveal the diversity and dynamics of BCR repertoires in unprecedented detail. Although this new type of sequence data has the potential to revolutionize our understanding of infection dynamics, quantitative analysis is complicated by the unique biology and high diversity of BCR sequences. Models and concepts from molecular evolution and phylogenetics that have been applied successfully to rapidly evolving pathogen populations are increasingly being adopted to study BCR diversity and divergence within individuals. However, BCR dynamics may violate key assumptions of many standard evolutionary methods, as they do not descend from a single ancestor, and experience biased mutation. Here, we review the application of evolutionary models to BCR repertoires and discuss the issues we believe need be addressed for this interdisciplinary field to flourish.
Journal Article
Repertoire-wide phylogenetic models of B cell molecular evolution reveal evolutionary signatures of aging and vaccination
2019
In order to produce effective antibodies, B cells undergo rapid somatic hypermutation (SHM) and selection for binding affinity to antigen via a process called affinity maturation. The similarities between this process and evolution by natural selection have led many groups to use phylogenetic methods to characterize the development of immunological memory, vaccination, and other processes that depend on affinity maturation. However, these applications are limited by the fact that most phylogenetic models are designed to be applied to individual lineages comprising genetically diverse sequences, while B cell repertoires often consist of hundreds to thousands of separate low-diversity lineages. Further, several features of affinity maturation violate important assumptions in standard phylogenetic models. Here, we introduce a hierarchical phylogenetic framework that integrates information from all lineages in a repertoire to more precisely estimate model parameters while simultaneously incorporating the unique features of SHM. We demonstrate the power of this repertoire-wide approach by characterizing previously undescribed phenomena in affinity maturation. First, we find evidence consistent with age-related changes in SHM hot-spot targeting. Second, we identify a consistent relationship between increased tree length and signs of increased negative selection, apparent in the repertoires of recently vaccinated subjects and those without any known recent infections or vaccinations. This suggests that B cell lineages shift toward negative selection over time as a general feature of affinity maturation. Our study provides a framework for undertaking repertoire-wide phylogenetic testing of SHM hypotheses and provides a means of characterizing dynamics of mutation and selection during affinity maturation.
Journal Article
8.2% of the Human Genome Is Constrained: Variation in Rates of Turnover across Functional Element Classes in the Human Lineage
2014
Ten years on from the finishing of the human reference genome sequence, it remains unclear what fraction of the human genome confers function, where this sequence resides, and how much is shared with other mammalian species. When addressing these questions, functional sequence has often been equated with pan-mammalian conserved sequence. However, functional elements that are short-lived, including those contributing to species-specific biology, will not leave a footprint of long-lasting negative selection. Here, we address these issues by identifying and characterising sequence that has been constrained with respect to insertions and deletions for pairs of eutherian genomes over a range of divergences. Within noncoding sequence, we find increasing amounts of mutually constrained sequence as species pairs become more closely related, indicating that noncoding constrained sequence turns over rapidly. We estimate that half of present-day noncoding constrained sequence has been gained or lost in approximately the last 130 million years (half-life in units of divergence time, d1/2 = 0.25-0.31). While enriched with ENCODE biochemical annotations, much of the short-lived constrained sequences we identify are not detected by models optimized for wider pan-mammalian conservation. Constrained DNase 1 hypersensitivity sites, promoters and untranslated regions have been more evolutionarily stable than long noncoding RNA loci which have turned over especially rapidly. By contrast, protein coding sequence has been highly stable, with an estimated half-life of over a billion years (d1/2 = 2.1-5.0). From extrapolations we estimate that 8.2% (7.1-9.2%) of the human genome is presently subject to negative selection and thus is likely to be functional, while only 2.2% has maintained constraint in both human and mouse since these species diverged. These results reveal that the evolutionary history of the human genome has been highly dynamic, particularly for its noncoding yet biologically functional fraction.
Journal Article
Lifestyle Patterns and Incidence of Cardiovascular Diseases, Cancer, Respiratory Diseases, and Type 2 Diabetes: A Large-Scale Prospective Cohort Study
2025
: Lifestyle factors often interact in complex ways when influencing chronic disease risk. We aimed to examine the prospective associations between empirically derived real-life lifestyle patterns (LPs) and the incidence of major chronic diseases, and to explore the linearity of the relationships between lifestyle summation scores and disease risk.
: We included adults free of cardiovascular diseases (CVDs), cancer, chronic respiratory diseases (CRDs), or type 2 diabetes (T2D) at baseline (2006-2013) from the Dutch Lifelines cohort. LPs and lifestyle summation scores were derived from baseline self-reported data on diet, physical activity, substance use, sleep, stress, and social connectedness, each categorised as healthy, moderately healthy, or unhealthy. Fine-Gray sub-hazard regression models assessed associations between LPs and disease incidence, with natural spline functions used to evaluate linearity in summation scores.
: Among 114,919 T2D-free, 131,248 cancer-free, 91,777 CRD-free, and 77,645 CVD-free participants, we observed 3114 T2D, 4685 cancer, 4133 CRDs, and 2850 CVD incident cases (median follow-up time: 8 years). Compared to the \"Unhealthy\" pattern, both the \"Healthy-in-a-balanced-way\" and \"Healthy-but-physically-inactive\" patterns were broadly significantly protective. The \"Unhealthy-but-no-substance-use\" pattern was associated with increased T2D risk (Sub-Hazard Ratio (SHR) = 1.27, 95% Confidence Interval (CI): 1.11-1.47) but reduced cancer risk (SHR = 0.85, 95%CI: 0.74-0.97). The \"Unhealthy-but-light-drinking-and-never-smoked\" pattern was protective for T2D (SHR = 0.89, 95%CI: 0.79-0.99). Linear associations were observed between lifestyle summation scores and disease risk, except for \"healthy lifestyle\" scores with T2D and \"unhealthy lifestyle\" scores with CRDs (non-linear
-value < 0.05).
: There are potential protective effects of healthy lifestyles on T2D, cancer, CRDs, and CVDs. However, the \"Unhealthy but no substance use\" demonstrated increased risk on T2D, protective effect on cancer and no significant effect on CRDs or CVDs. The relationship between combined lifestyle factors and NCD risk is complex and partly non-linear, showing diminishing benefits beyond certain thresholds, especially T2D and CRDs.
Journal Article
Genomic and Transcriptional Co-Localization of Protein-Coding and Long Non-Coding RNA Pairs in the Developing Brain
by
Ponting, Chris P.
,
Oliver, Peter L.
,
Ponjavic, Jasmina
in
Animals
,
Biological Transport
,
Brain - growth & development
2009
Besides protein-coding mRNAs, eukaryotic transcriptomes include many long non-protein-coding RNAs (ncRNAs) of unknown function that are transcribed away from protein-coding loci. Here, we have identified 659 intergenic long ncRNAs whose genomic sequences individually exhibit evolutionary constraint, a hallmark of functionality. Of this set, those expressed in the brain are more frequently conserved and are significantly enriched with predicted RNA secondary structures. Furthermore, brain-expressed long ncRNAs are preferentially located adjacent to protein-coding genes that are (1) also expressed in the brain and (2) involved in transcriptional regulation or in nervous system development. This led us to the hypothesis that spatiotemporal co-expression of ncRNAs and nearby protein-coding genes represents a general phenomenon, a prediction that was confirmed subsequently by in situ hybridisation in developing and adult mouse brain. We provide the full set of constrained long ncRNAs as an important experimental resource and present, for the first time, substantive and predictive criteria for prioritising long ncRNA and mRNA transcript pairs when investigating their biological functions and contributions to development and disease.
Journal Article