Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
17
result(s) for
"Kang, Dongwan D."
Sort by:
MetaBAT 2: an adaptive binning algorithm for robust and efficient genome reconstruction from metagenome assemblies
by
Kang, Dongwan D.
,
An, Hong
,
Li, Feng
in
Algorithms
,
BASIC BIOLOGICAL SCIENCES
,
Bioinformatics
2019
We previously reported on MetaBAT, an automated metagenome binning software tool to reconstruct single genomes from microbial communities for subsequent analyses of uncultivated microbial species. MetaBAT has become one of the most popular binning tools largely due to its computational efficiency and ease of use, especially in binning experiments with a large number of samples and a large assembly. MetaBAT requires users to choose parameters to fine-tune its sensitivity and specificity. If those parameters are not chosen properly, binning accuracy can suffer, especially on assemblies of poor quality. Here, we developed MetaBAT 2 to overcome this problem. MetaBAT 2 uses a new adaptive binning algorithm to eliminate manual parameter tuning. We also performed extensive software engineering optimization to increase both computational and memory efficiency. Comparing MetaBAT 2 to alternative software tools on over 100 real world metagenome assemblies shows superior accuracy and computing speed. Binning a typical metagenome assembly takes only a few minutes on a single commodity workstation. We therefore recommend the community adopts MetaBAT 2 for their metagenome binning experiments. MetaBAT 2 is open source software and available at https://bitbucket.org/berkeleylab/metabat .
Journal Article
MetaBAT, an efficient tool for accurately reconstructing single genomes from complex microbial communities
2015
Grouping large genomic fragments assembled from shotgun metagenomic sequences to deconvolute complex microbial communities, or metagenome binning, enables the study of individual organisms and their interactions. Because of the complex nature of these communities, existing metagenome binning methods often miss a large number of microbial species. In addition, most of the tools are not scalable to large datasets. Here we introduce automated software called MetaBAT that integrates empirical probabilistic distances of genome abundance and tetranucleotide frequency for accurate metagenome binning. MetaBAT outperforms alternative methods in accuracy and computational efficiency on both synthetic and real metagenome datasets. It automatically forms hundreds of high quality genome bins on a very large assembly consisting millions of contigs in a matter of hours on a single node. MetaBAT is open source software and available at https://bitbucket.org/berkeleylab/metabat.
Journal Article
The Epigenomic Landscape of Prokaryotes
by
Kang, Dongwan D.
,
Posfai, Janos
,
Morgan, Richard D.
in
60 APPLIED LIFE SCIENCES
,
BASIC BIOLOGICAL SCIENCES
,
Binding sites
2016
DNA methylation acts in concert with restriction enzymes to protect the integrity of prokaryotic genomes. Studies in a limited number of organisms suggest that methylation also contributes to prokaryotic genome regulation, but the prevalence and properties of such non-restriction-associated methylation systems remain poorly understood. Here, we used single molecule, real-time sequencing to map DNA modifications including m6A, m4C, and m5C across the genomes of 230 diverse bacterial and archaeal species. We observed DNA methylation in nearly all (93%) organisms examined, and identified a total of 834 distinct reproducibly methylated motifs. This data enabled annotation of the DNA binding specificities of 620 DNA Methyltransferases (MTases), doubling known specificities for previously hard to study Type I, IIG and III MTases, and revealing their extraordinary diversity. Strikingly, 48% of organisms harbor active Type II MTases with no apparent cognate restriction enzyme. These active 'orphan' MTases are present in diverse bacterial and archaeal phyla and show motif specificities and methylation patterns consistent with functions in gene regulation and DNA replication. Our results reveal the pervasive presence of DNA methylation throughout the prokaryotic kingdoms, as well as the diversity of sequence specificities and potential functions of DNA methylation systems.
Journal Article
Profiling Early Lung Immune Responses in the Mouse Model of Tuberculosis
2011
Tuberculosis (TB) is caused by the intracellular bacteria Mycobacterium tuberculosis, and kills more than 1.5 million people every year worldwide. Immunity to TB is associated with the accumulation of IFNγ-producing T helper cell type 1 (Th1) in the lungs, activation of M.tuberculosis-infected macrophages and control of bacterial growth. However, very little is known regarding the early immune responses that mediate accumulation of activated Th1 cells in the M.tuberculosis-infected lungs. To define the induction of early immune mediators in the M.tuberculosis-infected lung, we performed mRNA profiling studies and characterized immune cells in M.tuberculosis-infected lungs at early stages of infection in the mouse model. Our data show that induction of mRNAs involved in the recognition of pathogens, expression of inflammatory cytokines, activation of APCs and generation of Th1 responses occurs between day 15 and day 21 post infection. The induction of these mRNAs coincides with cellular accumulation of Th1 cells and activation of myeloid cells in M.tuberculosis-infected lungs. Strikingly, we show the induction of mRNAs associated with Gr1+ cells, namely neutrophils and inflammatory monocytes, takes place on day 12 and coincides with cellular accumulation of Gr1+ cells in M.tuberculosis-infected lungs. Interestingly, in vivo depletion of Gr1+ neutrophils between days 10-15 results in decreased accumulation of Th1 cells on day 21 in M.tuberculosis-infected lungs without impacting overall protective outcomes. These data suggest that the recruitment of Gr1+ neutrophils is an early event that leads to production of chemokines that regulate the accumulation of Th1 cells in the M.tuberculosis-infected lungs.
Journal Article
Helminth-induced arginase-1 exacerbates lung inflammation and disease severity in tuberculosis
by
Kang, Dongwan D.
,
Ahmed, Mushtaq
,
Rangel-Moreno, Javier
in
Animals
,
Antigens
,
Arginase - blood
2015
Parasitic helminth worms, such as Schistosoma mansoni, are endemic in regions with a high prevalence of tuberculosis (TB) among the population. Human studies suggest that helminth coinfections contribute to increased TB susceptibility and increased rates of TB reactivation. Prevailing models suggest that T helper type 2 (Th2) responses induced by helminth infection impair Th1 immune responses and thereby limit Mycobacterium tuberculosis (Mtb) control. Using a pulmonary mouse model of Mtb infection, we demonstrated that S. mansoni coinfection or immunization with S. mansoni egg antigens can reversibly impair Mtb-specific T cell responses without affecting macrophage-mediated Mtb control. Instead, S. mansoni infection resulted in accumulation of high arginase-1-expressing macrophages in the lung, which formed type 2 granulomas and exacerbated inflammation in Mtb-infected mice. Treatment of coinfected animals with an antihelminthic improved Mtb-specific Th1 responses and reduced disease severity. In a genetically diverse mouse population infected with Mtb, enhanced arginase-1 activity was associated with increased lung inflammation. Moreover, in patients with pulmonary TB, lung damage correlated with increased serum activity of arginase-1, which was elevated in TB patients coinfected with helminths. Together, our data indicate that helminth coinfection induces arginase-1-expressing type 2 granulomas, thereby increasing inflammation and TB disease severity. These results also provide insight into the mechanisms by which helminth coinfections drive increased susceptibility, disease progression, and severity in TB.
Journal Article
Missing value imputation in high-dimensional phenomic data: imputable or not, and how?
2014
Background
In modern biomedical research of complex diseases, a large number of demographic and clinical variables, herein called phenomic data, are often collected and missing values (MVs) are inevitable in the data collection process. Since many downstream statistical and bioinformatics methods require complete data matrix, imputation is a common and practical solution. In high-throughput experiments such as microarray experiments, continuous intensities are measured and many mature missing value imputation methods have been developed and widely applied. Numerous methods for missing data imputation of microarray data have been developed. Large phenomic data, however, contain continuous, nominal, binary and ordinal data types, which void application of most methods. Though several methods have been developed in the past few years, not a single complete guideline is proposed with respect to phenomic missing data imputation.
Results
In this paper, we investigated existing imputation methods for phenomic data, proposed a self-training selection (STS) scheme to select the best imputation method and provide a practical guideline for general applications. We introduced a novel concept of \"imputability measure\" (IM) to identify missing values that are fundamentally inadequate to impute. In addition, we also developed four variations of K-nearest-neighbor (KNN) methods and compared with two existing methods, multivariate imputation by chained equations (MICE) and missForest. The four variations are imputation by variables (KNN-V), by subjects (KNN-S), their weighted hybrid (KNN-H) and an adaptively weighted hybrid (KNN-A). We performed simulations and applied different imputation methods and the STS scheme to three lung disease phenomic datasets to evaluate the methods. An R package \"phenomeImpute\" is made publicly available.
Conclusions
Simulations and applications to real datasets showed that MICE often did not perform well; KNN-A, KNN-H and random forest were among the top performers although no method universally performed the best. Imputation of missing values with low imputability measures increased imputation errors greatly and could potentially deteriorate downstream analyses. The STS scheme was accurate in selecting the optimal method by evaluating methods in a second layer of missingness simulation. All source files for the simulation and the real data analyses are available on the author's publication website.
Journal Article
Integrative phenotyping framework (iPF): integrative clustering of multiple omics data identifies novel lung disease subphenotypes
by
Kim, SungHwan
,
Kang, Dongwan D.
,
Tseng, George C.
in
Algorithms
,
Animal Genetics and Genomics
,
Biomedical and Life Sciences
2015
Background
The increased multi-omics information on carefully phenotyped patients in studies of complex diseases requires novel methods for data integration. Unlike continuous intensity measurements from most omics data sets, phenome data contain clinical variables that are binary, ordinal and categorical.
Results
In this paper we introduce an integrative phenotyping framework (iPF) for disease subtype discovery. A feature topology plot was developed for effective dimension reduction and visualization of multi-omics data. The approach is free of model assumption and robust to data noises or missingness. We developed a workflow to integrate homogeneous patient clustering from different omics data in an agglomerative manner and then visualized heterogeneous clustering of pairwise omics sources. We applied the framework to two batches of lung samples obtained from patients diagnosed with chronic obstructive lung disease (COPD) or interstitial lung disease (ILD) with well-characterized clinical (phenomic) data, mRNA and microRNA expression profiles. Application of iPF to the first training batch identified clusters of patients consisting of homogenous disease phenotypes as well as clusters with intermediate disease characteristics. Analysis of the second batch revealed a similar data structure, confirming the presence of intermediate clusters. Genes in the intermediate clusters were enriched with inflammatory and immune functional annotations, suggesting that they represent mechanistically distinct disease subphenotypes that may response to immunomodulatory therapies. The iPF software package and all source codes are publicly available.
Conclusions
Identification of subclusters with distinct clinical and biomolecular characteristics suggests that integration of phenomic and other omics information could lead to identification of novel mechanism-based disease sub-phenotypes.
Journal Article
The Epigenomic Landscape of Prokaryotes
by
Deutschbauer, Adam M
,
Posfai, Janos
,
Malmstrom, Rex R
in
Binding sites
,
DNA methylation
,
Genes
2016
DNA methylation acts in concert with restriction enzymes to protect the integrity of prokaryotic genomes. Studies in a limited number of organisms suggest that methylation also contributes to prokaryotic genome regulation, but the prevalence and properties of such non-restriction-associated methylation systems remain poorly understood. Here, we used single molecule, real-time sequencing to map DNA modifications including m6A, m4C, and m5C across the genomes of 230 diverse bacterial and archaeal species. We observed DNA methylation in nearly all (93%) organisms examined, and identified a total of 834 distinct reproducibly methylated motifs. This data enabled annotation of the DNA binding specificities of 620 DNA Methyltransferases (MTases), doubling known specificities for previously hard to study Type I, IIG and III MTases, and revealing their extraordinary diversity. Strikingly, 48% of organisms harbor active Type II MTases with no apparent cognate restriction enzyme. These active 'orphan' MTases are present in diverse bacterial and archaeal phyla and show motif specificities and methylation patterns consistent with functions in gene regulation and DNA replication. Our results reveal the pervasive presence of DNA methylation throughout the prokaryotic kingdoms, as well as the diversity of sequence specificities and potential functions of DNA methylation systems.
Journal Article
A robust statistical framework for reconstructing genomes from metagenomic data
2014
We present software that reconstructs genomes from shotgun metagenomic sequences using a reference-independent approach. This method permits the identification of OTUs in large complex communities where many species are unknown. Binning reduces the complexity of a metagenomic dataset enabling many downstream analyses previously unavailable. In this study we developed MetaBAT, a robust statistical framework that integrates probabilistic distances of genome abundance with sequence composition for automatic binning. Applying MetaBAT to a human gut microbiome dataset identified 173 highly specific genomes bins including many representing previously unidentified species.
Critical Assessment of Metagenome Interpretation—a benchmark of metagenomics software
2017
The Critical Assessment of Metagenome Interpretation (CAMI) community initiative presents results from its first challenge, a rigorous benchmarking of software for metagenome assembly, binning and taxonomic profiling.
Methods for assembly, taxonomic profiling and binning are key to interpreting metagenome data, but a lack of consensus about benchmarking complicates performance assessment. The Critical Assessment of Metagenome Interpretation (CAMI) challenge has engaged the global developer community to benchmark their programs on highly complex and realistic data sets, generated from ∼700 newly sequenced microorganisms and ∼600 novel viruses and plasmids and representing common experimental setups. Assembly and genome binning programs performed well for species represented by individual genomes but were substantially affected by the presence of related strains. Taxonomic profiling and binning programs were proficient at high taxonomic ranks, with a notable performance decrease below family level. Parameter settings markedly affected performance, underscoring their importance for program reproducibility. The CAMI results highlight current challenges but also provide a roadmap for software selection to answer specific research questions.
Journal Article