Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
2,459
result(s) for
"631/1647/514"
Sort by:
Machine learning random forest for predicting oncosomatic variant NGS analysis
by
Jacques, Coralie
,
Carlioz, Antoine
,
Beaufils, Nathalie
in
631/114
,
631/114/1305
,
631/114/1751
2021
Since 2017, we have used IonTorrent NGS platform in our hospital to diagnose and treat cancer. Analyzing variants at each run requires considerable time, and we are still struggling with some variants that appear correct on the metrics at first, but are found to be negative upon further investigation. Can any machine learning algorithm (ML) help us classify NGS variants? This has led us to investigate which ML can fit our NGS data and to develop a tool that can be routinely implemented to help biologists. Currently, one of the greatest challenges in medicine is processing a significant quantity of data. This is particularly true in molecular biology with the advantage of next-generation sequencing (NGS) for profiling and identifying molecular tumors and their treatment. In addition to bioinformatics pipelines, artificial intelligence (AI) can be valuable in helping to analyze mutation variants. Generating sequencing data from patient DNA samples has become easy to perform in clinical trials. However, analyzing the massive quantities of genomic or transcriptomic data and extracting the key biomarkers associated with a clinical response to a specific therapy requires a formidable combination of scientific expertise, biomolecular skills and a panel of bioinformatic and biostatistic tools, in which artificial intelligence is now successful in developing future routine diagnostics. However, cancer genome complexity and technical artifacts make identifying real variants challenging. We present a machine learning method for classifying pathogenic single nucleotide variants (SNVs), single nucleotide polymorphisms (SNPs), multiple nucleotide variants (MNVs), insertions, and deletions detected by NGS from different types of tumor specimens, such as: colorectal, melanoma, lung and glioma cancer. We compared our NGS data to different machine learning algorithms using the k-fold cross-validation method and to neural networks (deep learning) to measure the performance of the different ML algorithms and determine which one is a valid model for confirming NGS variant calls in cancer diagnosis. We trained our machine learning with 70% of our data samples, extracted from our local database (our data structure had 7 parameters: chromosome, position, exon, variant allele frequency, minor allele frequency, coverage and protein description) and validated it with the 30% remaining data. The model offering the best accuracy was chosen and implemented in the NGS analysis routine. Artificial intelligence was developed with the R script language version 3.6.0. We trained our model on 70% of 102,011 variants. Our best error rate (0.22%) was found with random forest machine learning (ntree = 500 and mtry = 4), with an AUC of 0.99. Neural networks achieved some good scores. The final trained model with the neural network achieved an accuracy of 98% and an ROC-AUC of 0.99 with validation data. We tested our RF model to interpret more than 2000 variants from our NGS database: 20 variants were misclassified (error rate < 1%). The errors were nomenclature problems and false positives. After adding false positives to our training database and implementing our RF model routinely, our error rate was always < 0.5%. The RF model shows excellent results for oncosomatic NGS interpretation and can easily be implemented in other molecular biology laboratories. AI is becoming increasingly important in molecular biomedical analysis and can be very helpful in processing medical data. Neural networks show a good capacity in variant classification, and in the future, they may be useful in predicting more complex variants.
Journal Article
Nanopore sequencing and assembly of a human genome with ultra-long reads
by
Quinlan, Aaron R
,
Richardson, Hollian
,
Olsen, Hugh E
in
45/23
,
631/1647/514/1948
,
631/1647/514/2254
2018
A human genome is sequenced and assembled
de novo
using a pocket-sized nanopore device.
We report the sequencing and assembly of a reference genome for the human GM12878 Utah/Ceph cell line using the MinION (Oxford Nanopore Technologies) nanopore sequencer. 91.2 Gb of sequence data, representing ∼30× theoretical coverage, were produced. Reference-based alignment enabled detection of large structural variants and epigenetic modifications.
De novo
assembly of nanopore reads alone yielded a contiguous assembly (NG50 ∼3 Mb). We developed a protocol to generate ultra-long reads (N50 > 100 kb, read lengths up to 882 kb). Incorporating an additional 5× coverage of these ultra-long reads more than doubled the assembly contiguity (NG50 ∼6.4 Mb). The final assembled genome was 2,867 million bases in size, covering 85.8% of the reference. Assembly accuracy, after incorporating complementary short-read sequencing data, exceeded 99.8%. Ultra-long reads enabled assembly and phasing of the 4-Mb major histocompatibility complex (MHC) locus in its entirety, measurement of telomere repeat length, and closure of gaps in the reference human genome assembly GRCh38.
Journal Article
Exponential scaling of single-cell RNA-seq in the past decade
by
Teichmann, Sarah A
,
Vento-tormo, Roser
,
Svensson, Valentine
in
Gene expression
,
Ribonucleic acid
,
Scaling
2018
Measurement of the transcriptomes of single cells has been feasible for only a few years, but it has become an extremely popular assay. While many types of analysis can be carried out and various questions can be answered by single-cell RNA-seq, a central focus is the ability to survey the diversity of cell types in a sample. Unbiased and reproducible cataloging of gene expression patterns in distinct cell types requires large numbers of cells. Technological developments and protocol improvements have fueled consistent and exponential increases in the number of cells that can be studied in single-cell RNA-seq analyses. In this Perspective, we highlight the key technological developments that have enabled this growth in the data obtained from single-cell RNA-seq experiments.
Journal Article
Exploring the hemicellulolytic properties and safety of Bacillus paralicheniformis as stepping stone in the use of new fibrolytic beneficial microbes
by
Maski, Soufiane
,
Oliveira Correia, Lydie
,
Architecture et fonction des macromolécules biologiques (AFMB)
in
631/1647
,
631/1647/2067
,
631/1647/2196
2023
Abstract Bacillus strains from the Moroccan Coordinated Collections of Microorganisms (CCMM) were characterised and tested for fibrolytic function and safety properties that would be beneficial for maintaining intestinal homeostasis, and recommend beneficial microbes in the field of health promotion research. Forty strains were investigated for their fibrolytic activities towards complex purified polysaccharides and natural fibres representative of dietary fibres (DFs) entering the colon for digestion. We demonstrated hemicellulolytic activities for nine strains of Bacillus aerius , re-identified as Bacillus paralicheniformis and Bacillus licheniformis , using xylan, xyloglucan or lichenan as purified polysaccharides, and orange, apple and carrot natural fibres, with strain- and substrate-dependent production of glycoside hydrolases (GHs). Our combined methods, based on enzymatic assays, secretome, and genome analyses, highlighted the hemicellulolytic activities of B. paralicheniformis and the secretion of specific glycoside hydrolases, in particular xylanases, compared to B. licheniformis . Genomic features of these strains revealed a complete set of GH genes dedicated to the degradation of various polysaccharides from DFs, including cellulose, hemicellulose and pectin, which may confer on the strains the ability to digest a variety of DFs. Preliminary experiments on the safety and immunomodulatory properties of B. paralicheniformis fibrolytic strains were evaluated in light of applications as beneficial microbes' candidates for health improvement. B. paralicheniformis CCMM B969 was therefore proposed as a new fibrolytic beneficial microbe candidate.
Journal Article
Integrated proteogenomic deep sequencing and analytics accurately identify non-canonical peptides in tumor immunopeptidomes
2020
Efforts to precisely identify tumor human leukocyte antigen (HLA) bound peptides capable of mediating T cell-based tumor rejection still face important challenges. Recent studies suggest that non-canonical tumor-specific HLA peptides derived from annotated non-coding regions could elicit anti-tumor immune responses. However, sensitive and accurate mass spectrometry (MS)-based proteogenomics approaches are required to robustly identify these non-canonical peptides. We present an MS-based analytical approach that characterizes the non-canonical tumor HLA peptide repertoire, by incorporating whole exome sequencing, bulk and single-cell transcriptomics, ribosome profiling, and two MS/MS search tools in combination. This approach results in the accurate identification of hundreds of shared and tumor-specific non-canonical HLA peptides, including an immunogenic peptide derived from an open reading frame downstream of the melanoma stem cell marker gene
ABCB5
. These findings hold great promise for the discovery of previously unknown tumor antigens for cancer immunotherapy.
Non-canonical HLA-bound peptides from presumed non-coding regions are potential targets for cancer immunotherapy, but their discovery remains challenging. Here, the authors integrate exome sequencing, transcriptomics, ribosome profiling, and immunopeptidomics to identify tumor-specific non-canonical HLA-bound peptides.
Journal Article
Alterations in intestinal microbiota diversity, composition, and function in patients with sarcopenia
2021
16S rRNA sequencing of human fecal samples has been tremendously successful in identifying microbiome changes associated with both aging and disease. A number of studies have described microbial alterations corresponding to physical frailty and nursing home residence among aging individuals. A gut-muscle axis through which the microbiome influences skeletal muscle growth/function has been hypothesized. However, the microbiome has yet to be examined in sarcopenia. Here, we collected fecal samples of 60 healthy controls (CON) and 27 sarcopenic (Case)/possibly sarcopenic (preCase) individuals and analyzed the intestinal microbiota using 16S rRNA sequencing. We observed an overall reduction in microbial diversity in Case and preCase samples. The genera
Lachnospira
,
Fusicantenibacter
,
Roseburia
,
Eubacterium
, and
Lachnoclostridium
—known butyrate producers—were significantly less abundant in Case and preCase subjects while
Lactobacillus
was more abundant. Functional pathways underrepresented in Case subjects included numerous transporters and phenylalanine, tyrosine, and tryptophan biosynthesis suggesting that protein processing and nutrient transport may be impaired. In contrast, lipopolysaccharide biosynthesis was overrepresented in Case and PreCase subjects suggesting that sarcopenia is associated with a pro-inflammatory metagenome. These analyses demonstrate structural and functional alterations in the intestinal microbiota that may contribute to loss of skeletal muscle mass and function in sarcopenia.
Journal Article
Long-read sequencing in the era of epigenomics and epitranscriptomics
2023
As long-read sequencing technologies continue to advance, the possibility of obtaining maps of DNA and RNA modifications at single-molecule resolution has become a reality. Here we highlight the opportunities and challenges posed by the use of long-read sequencing technologies to study epigenetic and epitranscriptomic marks and how this will affect the way in which we approach the study of health and disease states.
Journal Article
Nanopore long-read RNAseq reveals widespread transcriptional variation among the surface receptors of individual B cells
by
Byrne, Ashley
,
Olsen, Hugh E.
,
Vollmers, Christopher
in
45/91
,
631/1647/514/1949
,
631/1647/514/2254
2017
Understanding gene regulation and function requires a genome-wide method capable of capturing both gene expression levels and isoform diversity at the single-cell level. Short-read RNAseq is limited in its ability to resolve complex isoforms because it fails to sequence full-length cDNA copies of RNA molecules. Here, we investigate whether RNAseq using the long-read single-molecule Oxford Nanopore MinION sequencer is able to identify and quantify complex isoforms without sacrificing accurate gene expression quantification. After benchmarking our approach, we analyse individual murine B1a cells using a custom multiplexing strategy. We identify thousands of unannotated transcription start and end sites, as well as hundreds of alternative splicing events in these B1a cells. We also identify hundreds of genes expressed across B1a cells that display multiple complex isoforms, including several B cell-specific surface receptors. Our results show that we can identify and quantify complex isoforms at the single cell level.
Short-read RNA-seq is limited in its ability to resolve complex transcript isoforms since it cannot sequence full-length cDNA. Here the authors use Oxford Nanopore MinION and their Mandalorion analysis pipeline to measure complex isoforms in B1a cells.
Journal Article
Comparative transcriptomics in human and mouse
by
Breschi, Alessandra
,
Gingeras, Thomas R.
,
Guigó, Roderic
in
631/1647/334/1874/345
,
631/1647/514/1949
,
631/1647/514/2254
2017
Key Points
The mouse is the most widely used model organism to study human disease, but often mouse biology cannot be extrapolated to humans. A deep comparison of mouse and human physiology at the molecular level is essential for understanding under which circumstances the mouse can be a suitable model of human biology and for creating better mouse models. Advances in next-generation sequencing technologies fostered genome-wide annotation of functional DNA elements, enabling extensive comparison of the human and mouse genomes.
At the transcriptional level, human and mouse gene expression profiles are conserved overall, although the degree of conservation varies depending on the tissues and the genes that are compared. Therefore, the question of whether the human and mouse transcriptomes cluster preferentially by tissue or organ or by species does not have an answer overall, and it depends specifically on the genes being considered.
Conservation of expression is not a direct consequence of conservation in regulatory sequences, including promoters and enhancers. Although gene regulatory networks are preserved overall between human and mouse, transcription binding sites are often not conserved.
Inter-individual genetic variation can affect human gene expression, but such variation cannot be modelled in inbred strains of laboratory mice because their genetic variation is small compared to the human population. An expansion of the current studies on the relationship between genetic variation and gene expression in outbred mice might provide helpful insights to understand the same relationship in humans.
Emerging technologies — such single-cell genomics and single-cell spatial transcriptomics — and time series experiments will improve the annotation of human and mouse genomes, refine the current definitions of homologous cell types and homologous (molecular) phenotypes, and ultimately help scientists to identify which mouse models are the most appropriate to address a given biological question.
Next-generation sequencing technologies have enabled the comprehensive characterization of human and mouse genomes, including at the transcriptional level. This article reviews the degree of conservation of human and mouse transcriptomes, along with the challenges of identifying when the mouse is a suitable model of human physiology.
Cross-species comparisons of genomes, transcriptomes and gene regulation are now feasible at unprecedented resolution and throughput, enabling the comparison of human and mouse biology at the molecular level. Insights have been gained into the degree of conservation between human and mouse at the level of not only gene expression but also epigenetics and inter-individual variation. However, a number of limitations exist, including incomplete transcriptome characterization and difficulties in identifying orthologous phenotypes and cell types, which are beginning to be addressed by emerging technologies. Ultimately, these comparisons will help to identify the conditions under which the mouse is a suitable model of human physiology and disease, and optimize the use of animal models.
Journal Article
Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and Cufflinks
2012
Recent advances in high-throughput cDNA sequencing (RNA-seq) can reveal new genes and splice variants and quantify expression genome-wide in a single assay. The volume and complexity of data from RNA-seq experiments necessitate scalable, fast and mathematically principled analysis software. TopHat and Cufflinks are free, open-source software tools for gene discovery and comprehensive expression analysis of high-throughput mRNA sequencing (RNA-seq) data. Together, they allow biologists to identify new genes and new splice variants of known ones, as well as compare gene and transcript expression under two or more conditions. This protocol describes in detail how to use TopHat and Cufflinks to perform such analyses. It also covers several accessory tools and utilities that aid in managing data, including CummeRbund, a tool for visualizing RNA-seq analysis results. Although the procedure assumes basic informatics skills, these tools assume little to no background with RNA-seq analysis and are meant for novices and experts alike. The protocol begins with raw sequencing reads and produces a transcriptome assembly, lists of differentially expressed and regulated genes and transcripts, and publication-quality visualizations of analysis results. The protocol's execution time depends on the volume of transcriptome sequencing data and available computing resources but takes less than 1 d of computer time for typical experiments and ∼1 h of hands-on time.
Journal Article