Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
      More Filters
      Clear All
      More Filters
      Source
    • Language
707 result(s) for "Karp, D."
Sort by:
Pathway size matters: the influence of pathway granularity on over-representation (enrichment analysis) statistics
Background Enrichment or over-representation analysis is a common method used in bioinformatics studies of transcriptomics, metabolomics, and microbiome datasets. The key idea behind enrichment analysis is: given a set of significantly expressed genes (or metabolites), use that set to infer a smaller set of perturbed biological pathways or processes, in which those genes (or metabolites) play a role. Enrichment computations rely on collections of defined biological pathways and/or processes, which are usually drawn from pathway databases. Although practitioners of enrichment analysis take great care to employ statistical corrections (e.g., for multiple testing), they appear unaware that enrichment results are quite sensitive to the pathway definitions that the calculation uses. Results We show that alternative pathway definitions can alter enrichment p -values by up to nine orders of magnitude, whereas statistical corrections typically alter enrichment p -values by only two orders of magnitude. We present multiple examples where the smaller pathway definitions used in the EcoCyc database produces stronger enrichment p -values than the much larger pathway definitions used in the KEGG database; we demonstrate that to attain a given enrichment p -value, KEGG-based enrichment analyses require 1.3–2.0 times as many significantly expressed genes as does EcoCyc-based enrichment analyses. The large pathways in KEGG are problematic for another reason: they blur together multiple (as many as 21) biological processes. When such a KEGG pathway receives a high enrichment p -value, which of its component processes is perturbed is unclear, and thus the biological conclusions drawn from enrichment of large pathways are also in question. Conclusions The choice of pathway database used in enrichment analyses can have a much stronger effect on the enrichment results than the statistical corrections used in these analyses.
The BioCyc Metabolic Network Explorer
Background The Metabolic Network Explorer is a new addition to the BioCyc.org website and the Pathway Tools software suite that supports the interactive exploration of metabolic networks. Any metabolic network visualization tool must by necessity show only a subset of all possible metabolite connections, or the results will be visually overwhelming. Existing tools, even those that purport to show an organism’s full metabolic network, limit the set of displayed connections based on predefined pathways or other preselected criteria. We sought instead to provide a tool that would give the user dynamic control over which connections to follow. Results The Metabolic Network Explorer is an easy-to-use, web-based software tool that allows the user to specify a starting metabolite of interest and interactively explore its immediate metabolic neighborhood in either or both directions to any desired depth, letting the user select from the full set of connected reactions. Although, as for other tools, only a small portion of the metabolic network is visible at a time, that portion is selected by the user, based on the full reaction complement, and it is easy to switch among alternate paths of interest. The display is intuitive, customizable, and provides copious links to more detailed information pages. Conclusions The Metabolic Network Explorer fills a gap in the set of metabolic network visualization tools and complements other modes of exploration. Its primary strengths are its ease of use, diagrams that are intuitive to biologists, and its integration with the broader corpus of data provided by a BioCyc Pathway/Genome Database.
A systematic comparison of the MetaCyc and KEGG pathway databases
Background The MetaCyc and KEGG projects have developed large metabolic pathway databases that are used for a variety of applications including genome analysis and metabolic engineering. We present a comparison of the compound, reaction, and pathway content of MetaCyc version 16.0 and a KEGG version downloaded on Feb-27-2012 to increase understanding of their relative sizes, their degree of overlap, and their scope. To assess their overlap, we must know the correspondences between compounds, reactions, and pathways in MetaCyc, and those in KEGG. We devoted significant effort to computational and manual matching of these entities, and we evaluated the accuracy of the correspondences. Results KEGG contains 179 module pathways versus 1,846 base pathways in MetaCyc; KEGG contains 237 map pathways versus 296 super pathways in MetaCyc. KEGG pathways contain 3.3 times as many reactions on average as do MetaCyc pathways, and the databases employ different conceptualizations of metabolic pathways. KEGG contains 8,692 reactions versus 10,262 for MetaCyc. 6,174 KEGG reactions are components of KEGG pathways versus 6,348 for MetaCyc. KEGG contains 16,586 compounds versus 11,991 for MetaCyc. 6,912 KEGG compounds act as substrates in KEGG reactions versus 8,891 for MetaCyc. MetaCyc contains a broader set of database attributes than does KEGG, such as relationships from a compound to enzymes that it regulates, identification of spontaneous reactions, and the expected taxonomic range of metabolic pathways. MetaCyc contains many pathways not found in KEGG, from plants, fungi, metazoa, and actinobacteria; KEGG contains pathways not found in MetaCyc, for xenobiotic degradation, glycan metabolism, and metabolism of terpenoids and polyketides. MetaCyc contains fewer unbalanced reactions, which facilitates metabolic modeling such as using flux-balance analysis. MetaCyc includes generic reactions that may be instantiated computationally. Conclusions KEGG contains significantly more compounds than does MetaCyc, whereas MetaCyc contains significantly more reactions and pathways than does KEGG, in particular KEGG modules are quite incomplete. The number of reactions occurring in pathways in the two DBs are quite similar.
AB1044 POSITIVE ANA TESTING IN AN ACADEMIC MEDICAL CENTER: IMPACT ON DIAGNOSIS
Background:sAnti-nuclear antibodies (ANA) are common in systemic rheumatic diseases, non-rheumatic autoimmune diseases and the general population. In clinical practice, testing for ANA can inform further clinical diagnoses even in the absence of symptoms to suggest SLE or connective tissue disease. This study was undertaken to evaluate whether ANA testing result informed clinical diagnosis.Objectives:1). Determine the frequency of ANA testing in an academic medical center, 2). Determine the demographics and principal diagnoses of people undergoing ANA testing in an academic medical center, and 3) Determine the impact of positive ANA determinations on diagnosis and patient care.Methods:The UT Southwestern Medical Center IRB approved this study. Data were obtained by SQL queries of the Epic electronic health record. The study population included all patients for whom an ANA was ordered at UT Southwestern Medical Center January 1, 2010 to June 30, 2017. The titer and pattern of the ANA as well as the results of ENA or anti-dsDNA testing were documented. Patient characteristics included age and sex. Encounter characteristics included the date of testing, frequency of testing, primary encounter diagnosis, and provider specialty. Chart review was performed by a board-certified rheumatologist in patients who had a change in diagnosis after a positive ANA.Results:During the study period, a total of 33,270 ANA tests were ordered in 28,659 unique patients. 22,529 of the ANAs were tested in outpatients, representing 0.7% of all office visits during this time. Forty-nine percent of the ANA tests were positive at a titer of 1:80 or greater; slightly more women (51%) had a positive ANA (≥1:80) than did men (43%). 54.5% of the positive ANA in women were 1:320 or greater vs. 41.6% in men (p<0.0001). In 118 patients with a positive ANA, ICD 9 coding changed from a non-rheumatic disease diagnosis to a rheumatic disease diagnosis after ANA testing. Chart review showed that 26 of these patients had testing for joint pain. In 7 the note described this as inflammatory, while in 19 it was not specified. In five of the seven patients with a prior diagnosis of rheumatoid arthritis, the diagnosis was changed to SLE, and one was changed to Sicca syndrome. The 19 with non-inflammatory arthritis were diagnosed UCTD (7), with sicca or Sjögren’s (6), drug induced SLE (1) and SLE (3) although none of these fulfilled the 2019 EULAR/ACR criteria. 16 patients had a prior diagnosis of SLE that was confirmed. 14 had a initial diagnosis of “positive ANA” and 13 had their diagnosis changed to a systemic rheumatologic disease (SLE, incomplete SLE, Sjögren’s syndrome, UCTD, or systemic sclerosis). 9 patients had autoimmune liver disease confirmed and 8 had confirmed dermatomyositis. In 1 of 12 patients that carried a diagnosis of rheumatologic disorder (MCTD or Sjögrens) the diagnosis was changed to SLE while the other 11 had the diagnosis stay the same. In most of the remaining patients with either symptoms specific to SLE (pleurisy, pancytopenia, or non-specific hearing loss), the diagnosis was changed to UCTD or sicca syndrome.Conclusion:ANA testing in the inpatient and outpatient setting is common. Diagnoses precipitating testing are most often non-rheumatic conditions. A positive ANA result changed the clinical diagnosis in a small percentage of patients and rarely informed treatment.REFERENCES:NIL.Acknowledgements:NIL.Disclosure of Interests:David Karp Support to university from UCB, Eli Lilly, Genentech, Bristol-Myers Squibb, Novartis, and Biogen., Bonnie Bermas: None declared, Shivani Kottur: None declared.
Machine learning methods for metabolic pathway prediction
Background A key challenge in systems biology is the reconstruction of an organism's metabolic network from its genome sequence. One strategy for addressing this problem is to predict which metabolic pathways, from a reference database of known pathways, are present in the organism, based on the annotated genome of the organism. Results To quantitatively validate methods for pathway prediction, we developed a large \"gold standard\" dataset of 5,610 pathway instances known to be present or absent in curated metabolic pathway databases for six organisms. We defined a collection of 123 pathway features, whose information content we evaluated with respect to the gold standard. Feature data were used as input to an extensive collection of machine learning (ML) methods, including naïve Bayes, decision trees, and logistic regression, together with feature selection and ensemble methods. We compared the ML methods to the previous PathoLogic algorithm for pathway prediction using the gold standard dataset. We found that ML-based prediction methods can match the performance of the PathoLogic algorithm. PathoLogic achieved an accuracy of 91% and an F-measure of 0.786. The ML-based prediction methods achieved accuracy as high as 91.2% and F-measure as high as 0.787. The ML-based methods output a probability for each predicted pathway, whereas PathoLogic does not, which provides more information to the user and facilitates filtering of predicted pathways. Conclusions ML methods for pathway prediction perform as well as existing methods, and have qualitative advantages in terms of extensibility, tunability, and explainability. More advanced prediction methods and/or more sophisticated input features may improve the performance of ML methods. However, pathway prediction performance appears to be limited largely by the ability to correctly match enzymes to the reactions they catalyze based on genome annotations.
Evaluation of reaction gap-filling accuracy by randomization
Background Completion of genome-scale flux-balance models using computational reaction gap-filling is a widely used approach, but its accuracy is not well known. Results We report on computational experiments of reaction gap filling in which we generated degraded versions of the EcoCyc-20.0-GEM model by randomly removing flux-carrying reactions from a growing model. We gap-filled the degraded models and compared the resulting gap-filled models with the original model. Gap-filling was performed by the Pathway Tools MetaFlux software using its General Development Mode (GenDev) and its Fast Development Mode (FastDev). We explored 12 GenDev variants including two linear solvers (SCIP and CPLEX) for solving the Mixed Integer Linear Programming (MILP) problems for gap filling; three different sets of linear constraints were applied; and two MILP methods were implemented. We compared these 13 variants according to accuracy, speed, and amount of information returned to the user. Conclusions We observed large variation among the performance of the 13 gap-filling variants. Although no variant was best in all dimensions, we found one variant that was fast, accurate, and returned more information to the user. Some gap-filling variants were inaccurate, producing solutions that were non-minimum or invalid (did not enable model growth). The best GenDev variant showed a best average precision of 87% and a best average recall of 61%. FastDev showed an average precision of 71% and an average recall of 59%. Thus, using the most accurate variant, approximately 13% of the gap-filled reactions were incorrect (were not the reactions removed from the model), and 39% of gap-filled reactions were not found, suggesting that curation is still an important aspect of metabolic-model development.
The MultiOmics Explainer: explaining omics results in the context of a pathway/genome database
Background High-throughput experiments can bring to light associations between genes, proteins and/or metabolites, many of which will be explainable by existing knowledge. Our aim is to speed elucidation of such explanations and, in some cases, find explanations that scientists might otherwise overlook. Results We describe the MultiOmics Explainer, a new tool within the Pathway Tools software suite that leverages what is known about an organism’s metabolic and regulatory network to suggest explanations for the results of omics experiments. Querying a database such as EcoCyc, the MultiOmics Explainer searches the organism’s network of metabolic reactions, transporters, cofactors, enzyme substrate-level activation and inhibition relationships, and transcriptional and translational regulation relationships to identify paths of influence among input genes, proteins and metabolites. Results are presented in a combined metabolic and regulatory diagram. We present several examples of explanations generated for associations found in the Escherichia coli literature. Conclusions The MultiOmics Explainer is a valuable tool that helps researchers understand and interpret the results of their omics experiments in the context of what is known about an organism’s metabolic and regulatory network. It showcases the rich set of computational inferences that can be drawn from a database such as EcoCyc that encodes a diverse range of biological interactions.
Gene Dispensability in Escherichia coli Grown in Thirty Different Carbon Environments
While there has been much study of bacterial gene dispensability, there is a lack of comprehensive genome-scale examinations of the impact of gene deletion on growth in different carbon sources. In this context, a lot can be learned from such experiments in the model microbe Escherichia coli where much is already understood and there are existing tools for the investigation of carbon metabolism and physiology (1). Gene deletion studies have practical potential in the field of antibiotic drug discovery where there is emerging interest in bacterial central metabolism as a target for new antibiotics (2). Furthermore, some carbon utilization pathways have been shown to be critical for initiating and maintaining infection for certain pathogens and sites of infection (3–5). Here, with the use of high-throughput solid medium phenotyping methods, we have generated kinetic growth measurements for 3,796 genes under 30 different carbon source conditions. This data set provides a foundation for research that will improve our understanding of genes with unknown function, aid in predicting potential antibiotic targets, validate and advance metabolic models, and help to develop our understanding of E. coli metabolism. Central metabolism is a topic that has been studied for decades, and yet, this process is still not fully understood in Escherichia coli , perhaps the most amenable and well-studied model organism in biology. To further our understanding, we used a high-throughput method to measure the growth kinetics of each of 3,796 E. coli single-gene deletion mutants in 30 different carbon sources. In total, there were 342 genes (9.01%) encompassing a breadth of biological functions that showed a growth phenotype on at least 1 carbon source, demonstrating that carbon metabolism is closely linked to a large number of processes in the cell. We identified 74 genes that showed low growth in 90% of conditions, defining a set of genes which are essential in nutrient-limited media, regardless of the carbon source. The data are compiled into a Web application, Carbon Phenotype Explorer (CarPE), to facilitate easy visualization of growth curves for each mutant strain in each carbon source. Our experimental data matched closely with the predictions from the EcoCyc metabolic model which uses flux balance analysis to predict growth phenotypes. From our comparisons to the model, we found that, unexpectedly, phosphoenolpyruvate carboxylase ( ppc ) was required for robust growth in most carbon sources other than most trichloroacetic acid (TCA) cycle intermediates. We also identified 51 poorly annotated genes that showed a low growth phenotype in at least 1 carbon source, which allowed us to form hypotheses about the functions of these genes. From this list, we further characterized the ydhC gene and demonstrated its role in adenosine efflux. IMPORTANCE While there has been much study of bacterial gene dispensability, there is a lack of comprehensive genome-scale examinations of the impact of gene deletion on growth in different carbon sources. In this context, a lot can be learned from such experiments in the model microbe Escherichia coli where much is already understood and there are existing tools for the investigation of carbon metabolism and physiology (1). Gene deletion studies have practical potential in the field of antibiotic drug discovery where there is emerging interest in bacterial central metabolism as a target for new antibiotics (2). Furthermore, some carbon utilization pathways have been shown to be critical for initiating and maintaining infection for certain pathogens and sites of infection (3–5). Here, with the use of high-throughput solid medium phenotyping methods, we have generated kinetic growth measurements for 3,796 genes under 30 different carbon source conditions. This data set provides a foundation for research that will improve our understanding of genes with unknown function, aid in predicting potential antibiotic targets, validate and advance metabolic models, and help to develop our understanding of E. coli metabolism.
The Omics Dashboard for Interactive Exploration of Metabolomics and Multi-Omics Data
The Omics Dashboard is a software tool for interactive exploration and analysis of metabolomics, transcriptomics, proteomics, and multi-omics datasets. Organized as a hierarchy of cellular systems, the Dashboard at its highest level contains graphical panels for the full range of cellular systems, including biosynthesis, energy metabolism, and response to stimulus. Thus, the Dashboard top level surveys the state of the cell across a broad range of key systems in a single screen. Each Dashboard panel contains a series of X–Y plots depicting the aggregated omics data values relevant to different subsystems of that panel, e.g., subsystems within the biosynthesis panel include amino acid biosynthesis, carbohydrate biosynthesis and cofactor biosynthesis. Users can interactively drill down to focus in on successively lower-level subsystems of interest. In this article, we present for the first time the metabolomics analysis capabilities of the Omics Dashboard, along with significant new extensions to better accommodate metabolomics datasets, enable analysis and visualization of multi-omics datasets, and provide new data-filtering options.