Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
      More Filters
      Clear All
      More Filters
      Source
    • Language
214 result(s) for "Johnson, W. Evan"
Sort by:
ComBat-seq: batch effect adjustment for RNA-seq count data
The benefit of integrating batches of genomic data to increase statistical power is often hindered by batch effects, or unwanted variation in data caused by differences in technical factors across batches. It is therefore critical to effectively address batch effects in genomic data to overcome these challenges. Many existing methods for batch effects adjustment assume the data follow a continuous, bell-shaped Gaussian distribution. However in RNA-seq studies the data are typically skewed, over-dispersed counts, so this assumption is not appropriate and may lead to erroneous results. Negative binomial regression models have been used previously to better capture the properties of counts. We developed a batch correction method, ComBat-seq, using a negative binomial regression model that retains the integer nature of count data in RNA-seq studies, making the batch adjusted data compatible with common differential expression software packages that require integer counts. We show in realistic simulations that the ComBat-seq adjusted data results in better statistical power and control of false positives in differential expression compared to data adjusted by the other available methods. We further demonstrated in a real data example that ComBat-seq successfully removes batch effects and recovers the biological signal in the data.
Decontamination of ambient RNA in single-cell RNA-seq with DecontX
Droplet-based microfluidic devices have become widely used to perform single-cell RNA sequencing (scRNA-seq). However, ambient RNA present in the cell suspension can be aberrantly counted along with a cell’s native mRNA and result in cross-contamination of transcripts between different cell populations. DecontX is a novel Bayesian method to estimate and remove contamination in individual cells. DecontX accurately predicts contamination levels in a mouse-human mixture dataset and removes aberrant expression of marker genes in PBMC datasets. We also compare the contamination levels between four different scRNA-seq protocols. Overall, DecontX can be incorporated into scRNA-seq workflows to improve downstream analyses.
Metagenomic profiling pipelines improve taxonomic classification for 16S amplicon sequencing data
Most experiments studying bacterial microbiomes rely on the PCR amplification of all or part of the gene for the 16S rRNA subunit, which serves as a biomarker for identifying and quantifying the various taxa present in a microbiome sample. Several computational methods exist for analyzing 16S amplicon sequencing. However, the most-used bioinformatics tools cannot produce high quality genus-level or species-level taxonomic calls and may underestimate the potential accuracy of these calls. We used 16S sequencing data from mock bacterial communities to evaluate the sensitivity and specificity of several bioinformatics pipelines and genomic reference libraries used for microbiome analyses, concentrating on measuring the accuracy of species-level taxonomic assignments of 16S amplicon reads. We evaluated the tools DADA2, QIIME 2, Mothur, PathoScope 2, and Kraken 2 in conjunction with reference libraries from Greengenes, SILVA, Kraken 2, and RefSeq. Profiling tools were compared using publicly available mock community data from several sources, comprising 136 samples with varied species richness and evenness, several different amplified regions within the 16S rRNA gene, and both DNA spike-ins and cDNA from collections of plated cells. PathoScope 2 and Kraken 2, both tools designed for whole-genome metagenomics, outperformed DADA2, QIIME 2 using the DADA2 plugin, and Mothur, which are theoretically specialized for 16S analyses. Evaluations of reference libraries identified the SILVA and RefSeq/Kraken 2 Standard libraries as superior in accuracy compared to Greengenes. These findings support PathoScope and Kraken 2 as fully capable, competitive options for genus- and species-level 16S amplicon sequencing data analysis, whole genome sequencing, and metagenomics data tools.
Comprehensive generation, visualization, and reporting of quality control metrics for single-cell RNA sequencing data
Single-cell RNA sequencing (scRNA-seq) can be used to gain insights into cellular heterogeneity within complex tissues. However, various technical artifacts can be present in scRNA-seq data and should be assessed before performing downstream analyses. While several tools have been developed to perform individual quality control (QC) tasks, they are scattered in different packages across several programming environments. Here, to streamline the process of generating and visualizing QC metrics for scRNA-seq data, we built the SCTK-QC pipeline within the singleCellTK R package. The SCTK-QC workflow can import data from several single-cell platforms and preprocessing tools and includes steps for empty droplet detection, generation of standard QC metrics, prediction of doublets, and estimation of ambient RNA. It can run on the command line, within the R console, on the cloud platform or with an interactive graphical user interface. Overall, the SCTK-QC pipeline streamlines and standardizes the process of performing QC for scRNA-seq data. Quality control (QC) is a crucial step in single-cell RNA-seq data analysis. Here, the authors present the SCTK-QC pipeline which generates and visualizes a comprehensive set of QC metrics to streamline the process of detecting and removing poor quality cells and other artifacts.
Inhibition of colony stimulating factor 1 receptor corrects maternal inflammation-induced microglial and synaptic dysfunction and behavioral abnormalities
Maternal immune activation (MIA) disrupts the central innate immune system during a critical neurodevelopmental period. Microglia are primary innate immune cells in the brain although their direct influence on the MIA phenotype is largely unknown. Here we show that MIA alters microglial gene expression with upregulation of cellular protrusion/neuritogenic pathways, concurrently causing repetitive behavior, social deficits, and synaptic dysfunction to layer V intrinsically bursting pyramidal neurons in the prefrontal cortex of mice. MIA increases plastic dendritic spines of the intrinsically bursting neurons and their interaction with hyper-ramified microglia. Treating MIA offspring by colony stimulating factor 1 receptor inhibitors induces depletion and repopulation of microglia, and corrects protein expression of the newly identified MIA-associated neuritogenic molecules in microglia, which coalesces with correction of MIA-associated synaptic, neurophysiological, and behavioral abnormalities. Our study demonstrates that maternal immune insults perturb microglial phenotypes and influence neuronal functions throughout adulthood, and reveals a potent effect of colony stimulating factor 1 receptor inhibitors on the correction of MIA-associated microglial, synaptic, and neurobehavioral dysfunctions.
Alternative empirical Bayes models for adjusting for batch effects in genomic studies
Background Combining genomic data sets from multiple studies is advantageous to increase statistical power in studies where logistical considerations restrict sample size or require the sequential generation of data. However, significant technical heterogeneity is commonly observed across multiple batches of data that are generated from different processing or reagent batches, experimenters, protocols, or profiling platforms. These so-called batch effects often confound true biological relationships in the data, reducing the power benefits of combining multiple batches, and may even lead to spurious results in some combined studies. Therefore there is significant need for effective methods and software tools that account for batch effects in high-throughput genomic studies. Results Here we contribute multiple methods and software tools for improved combination and analysis of data from multiple batches. In particular, we provide batch effect solutions for cases where the severity of the batch effects is not extreme, and for cases where one high-quality batch can serve as a reference, such as the training set in a biomarker study. We illustrate our approaches and software in both simulated and real data scenarios. Conclusions We demonstrate the value of these new contributions compared to currently established approaches in the specified batch correction situations.
Tackling the widespread and critical impact of batch effects in high-throughput data
Batch effects can lead to incorrect biological conclusions but are not widely considered. The authors show that batch effects are relevant to a range of high-throughput 'omics' data sets and are crucial to address. They also explain how batch effects can be mitigated. High-throughput technologies are widely used, for example to assay genetic variants, gene and protein expression, and epigenetic modifications. One often overlooked complication with such studies is batch effects, which occur because measurements are affected by laboratory conditions, reagent lots and personnel differences. This becomes a major problem when batch effects are correlated with an outcome of interest and lead to incorrect conclusions. Using both published studies and our own analyses, we argue that batch effects (as well as other technical and biological artefacts) are widespread and critical to address. We review experimental and computational approaches for doing so.
A comprehensive update to the Mycobacterium tuberculosis H37Rv reference genome
H37Rv is the most widely used Mycobacterium tuberculosis strain, and its genome is globally used as the M. tuberculosis reference sequence. Here, we present Bact-Builder, a pipeline that uses consensus building to generate complete and accurate bacterial genome sequences and apply it to three independently cultured and sequenced H37Rv aliquots of a single laboratory stock. Two of the 4,417,942 base-pair long H37Rv assemblies are 100% identical, with the third differing by a single nucleotide. Compared to the existing H37Rv reference, the new sequence contains ~6.4 kb additional base pairs, encoding ten new regions that include insertions in PE/PPE genes and new paralogs of esxN and esxJ , which are differentially expressed compared to the reference genes. New sequencing and de novo assemblies with Bact-Builder confirm that all 10 regions, plus small additional polymorphisms, are also present in the commonly used H37Rv strains NR123, TMC102, and H37Rv1998. Thus, Bact-Builder shows promise as an improved method to perform accurate and reproducible de novo assemblies of bacterial genomes, and our work provides important updates to the primary M. tuberculosis reference genome. H37Rv is the most widely used Mycobacterium tuberculosis strain, and its genome is the reference sequence for this pathogen. Here, Chitale et al. present a bioinformatic pipeline for accurate assembly of bacterial genome sequences, and use it to provide important updates to the M. tuberculosis reference genome.
Respiratory syncytial virus M2-1 protein associates non-specifically with viral messenger RNA and with specific cellular messenger RNA transcripts
Respiratory syncytial virus (RSV) is a major cause of respiratory disease in infants and the elderly. RSV is a non-segmented negative strand RNA virus. The viral M2-1 protein plays a key role in viral transcription, serving as an elongation factor to enable synthesis of full-length mRNAs. M2-1 contains an unusual CCCH zinc-finger motif that is conserved in the related human metapneumovirus M2-1 protein and filovirus VP30 proteins. Previous biochemical studies have suggested that RSV M2-1 might bind to specific virus RNA sequences, such as the transcription gene end signals or poly A tails, but there was no clear consensus on what RSV sequences it binds. To determine if M2-1 binds to specific RSV RNA sequences during infection, we mapped points of M2-1:RNA interactions in RSV-infected cells at 8 and 18 hours post infection using crosslinking immunoprecipitation with RNA sequencing (CLIP-Seq). This analysis revealed that M2-1 interacts specifically with positive sense RSV RNA, but not negative sense genome RNA. It also showed that M2-1 makes contacts along the length of each viral mRNA, indicating that M2-1 functions as a component of the transcriptase complex, transiently associating with nascent mRNA being extruded from the polymerase. In addition, we found that M2-1 binds specific cellular mRNAs. In contrast to the situation with RSV mRNA, M2-1 binds discrete sites within cellular mRNAs, with a preference for A/U rich sequences. These results suggest that in addition to its previously described role in transcription elongation, M2-1 might have an additional role involving cellular RNA interactions.
Bridging the gap between R and Python in bulk transcriptomic data analysis with InMoose
We introduce InMoose, an open-source Python environment aimed at omic data analysis. We illustrate its capabilities for bulk transcriptomic data analysis. Due to its wide adoption, Python has grown as a de facto standard in fields increasingly important for bioinformatic pipelines, such as data science, machine learning, or artificial intelligence (AI). As a general-purpose language, Python is also recognized for its versatility and scalability. InMoose aims at bringing state-of-the-art tools, historically written in R, to the Python ecosystem. InMoose focuses on providing drop-in replacements for R tools, to ensure consistency and reproducibility between R-based and Python-based pipelines. The first development phase has focused on bulk transcriptomic data, with current capabilities encompassing data simulation, batch effect correction, and differential analysis and meta-analysis.