Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
1,215
result(s) for
"Statistical workflow"
Sort by:
Visualization in Bayesian workflow
by
Gabry, Jonah
,
Betancourt, Michael
,
Gelman, Andrew
in
Bayesian analysis
,
Bayesian data analysis
,
Data analysis
2019
Bayesian data analysis is about more than just computing a posterior distribution, and Bayesian visualization is about more than trace plots of Markov chains. Practical Bayesian data analysis, like all data analysis, is an iterative process of model building, inference, model checking and evaluation, and model expansion. Visualization is helpful in each of these stages of the Bayesian workflow and it is indispensable when drawing inferences from the types of modern, high dimensional models that are used by applied researchers.
Journal Article
A new statistical workflow (R-packages based) to investigate associations between one variable of interest and the metabolome
by
Rist, Manuela J
,
Frommherz, Lara
,
Ferrario, Paola G
in
Datasets
,
Metabolomics
,
Statistical analysis
2023
IntroductionIn metabolomics, the investigation of associations between the metabolome and one trait of interest is a key research question. However, statistical analyses of such associations are often challenging. Statistical tools enabling resilient verification and clear presentation are therefore highly desired.ObjectivesOur aim is to provide a contribution for statistical analysis of metabolomics data, offering a widely applicable open-source statistical workflow, which considers the intrinsic complexity of metabolomics data.MethodsWe combined selected R packages tailored for all properties of heterogeneous metabolomics datasets, where metabolite parameters typically (i) are analyzed in different matrices, (ii) are measured on different analytical platforms with different precision, (iii) are analyzed by targeted as well as non-targeted methods, (iv) are scaled variously, (v) reveal heterogeneous variances, (vi) may be correlated, (vii) may have only few values or values below a detection limit, or (viii) may be incomplete.ResultsThe code is shared entirely and freely available. The workflow output is a table of metabolites associated with a trait of interest and a compact plot for high-quality results visualization. The workflow output and its utility are presented by applying it to two previously published datasets: one dataset from our own lab and another dataset taken from the repository MetaboLights.ConclusionRobustness and benefits of the statistical workflow were clearly demonstrated, and everyone can directly re-use it for analysis of own data.
Journal Article
Development of a Statistical Workflow for Screening Protein Extracts Based on Their Nutritional Composition and Digestibility: Application to Elderly
by
Ferraro, Vincenza
,
Ferreira, Claude De Oliviera
,
Sante-Lhoutellier, Véronique
in
Aged
,
Amino acids
,
Analysis
2020
The objective of the study is to develop a workflow to screen protein extracts and identify their nutritional potential as high quality nutritional culinary aids for recipes for the elderly. Twenty-seven protein extracts of animal, vegetable, and dairy origin were characterized. We studied their fate by monitoring static in vitro digestion, mimicking the physiological digestion conditions of the elderly. At the end of the gastric and intestinal phase, global measurements of digestibility and antioxidant bioactivities were performed. The statistical analysis workflow developed allowed: (i) synthesizing the compositional and nutritional information of each protein extract by creating latent variables, and (ii) comparing them. The links between variables and similarities between protein extracts were visualized using a heat map. A hierarchical cluster analysis allowed reducing the 48 quantitative variables into 15 qualitative latent variables (clusters). The application of the k-means method on each cluster enable to classify the protein extracts by level. This defined level was used as categorical value. Multiple correspondence analysis revealed groups of protein extracts with varied patterns. This workflow allowed the comparison/hierarchization between protein extracts and the creation of a tool to select the most interesting ones on the basis of their nutritional quality.
Journal Article
fMRIPrep: a robust preprocessing pipeline for functional MRI
by
Kent, James D
,
Esteban, Oscar
,
Durnez, Joke
in
Data collection
,
Data processing
,
Functional magnetic resonance imaging
2019
Preprocessing of functional magnetic resonance imaging (fMRI) involves numerous steps to clean and standardize the data before statistical analysis. Generally, researchers create ad hoc preprocessing workflows for each dataset, building upon a large inventory of available tools. The complexity of these workflows has snowballed with rapid advances in acquisition and processing. We introduce fMRIPrep, an analysis-agnostic tool that addresses the challenge of robust and reproducible preprocessing for fMRI data. fMRIPrep automatically adapts a best-in-breed workflow to the idiosyncrasies of virtually any dataset, ensuring high-quality preprocessing without manual intervention. By introducing visual assessment checkpoints into an iterative integration framework for software testing, we show that fMRIPrep robustly produces high-quality results on a diverse fMRI data collection. Additionally, fMRIPrep introduces less uncontrolled spatial smoothness than observed with commonly used preprocessing tools. fMRIPrep equips neuroscientists with an easy-to-use and transparent preprocessing workflow, which can help ensure the validity of inference and the interpretability of results.
Journal Article
Chemometric analysis in Raman spectroscopy from experimental design to machine learning–based modeling
by
Guo, Shuxia
,
Popp, Jürgen
,
Bocklitz, Thomas
in
631/114/2164
,
639/638/440/527/1821
,
639/705/1042
2021
Raman spectroscopy is increasingly being used in biology, forensics, diagnostics, pharmaceutics and food science applications. This growth is triggered not only by improvements in the computational and experimental setups but also by the development of chemometric techniques. Chemometric techniques are the analytical processes used to detect and extract information from subtle differences in Raman spectra obtained from related samples. This information could be used to find out, for example, whether a mixture of bacterial cells contains different species, or whether a mammalian cell is healthy or not. Chemometric techniques include spectral processing (ensuring that the spectra used for the subsequent computational processes are as clean as possible) as well as the statistical analysis of the data required for finding the spectral differences that are most useful for differentiation between, for example, different cell types. For Raman spectra, this analysis process is not yet standardized, and there are many confounding pitfalls. This protocol provides guidance on how to perform a Raman spectral analysis: how to avoid these pitfalls, and strategies to circumvent problematic issues. The protocol is divided into four parts: experimental design, data preprocessing, data learning and model transfer. We exemplify our workflow using three example datasets where the spectra from individual cells were collected in single-cell mode, and one dataset where the data were collected from a raster scanning–based Raman spectral imaging experiment of mice tissue. Our aim is to help move Raman-based technologies from proof-of-concept studies toward real-world applications.
Raman spectroscopy is increasingly being used in biological assays and studies. This protocol provides guidance for performing chemometric analysis to detect and extract information relating to the chemical differences between biological samples.
Journal Article
Probabilistic Integration
by
Oates, Chris J.
,
Osborne, Michael A.
,
Girolami, Mark
in
Computation
,
Computer graphics
,
Discretization
2019
A research frontier has emerged in scientific computation, wherein discretisation error is regarded as a source of epistemic uncertainty that can be modelled. This raises several statistical challenges, including the design of statistical methods that enable the coherent propagation of probabilities through a (possibly deterministic) computational work-flow, in order to assess the impact of discretisation error on the computer output. This paper examines the case for probabilistic numerical methods in routine statistical computation. Our focus is on numerical integration, where a probabilistic integrator is equipped with a full distribution over its output that reflects the fact that the integrand has been discretised. Our main technical contribution is to establish, for the first time, rates of posterior contraction for one such method. Several substantial applications are provided for illustration and critical evaluation, including examples from statistical modelling, computer graphics and a computer model for an oil reservoir.
Journal Article
Data processing, multi-omic pathway mapping, and metabolite activity analysis using XCMS Online
by
sberg, Erica M
,
Hilmers, Brian
,
Huan, Tao
in
Bioinformatics
,
Biological activity
,
Biological effects
2018
Systems biology is the study of complex living organisms, and as such, analysis on a systems-wide scale involves the collection of information-dense data sets that are representative of an entire phenotype. To uncover dynamic biological mechanisms, bioinformatics tools have become essential to facilitating data interpretation in large-scale analyses. Global metabolomics is one such method for performing systems biology, as metabolites represent the downstream functional products of ongoing biological processes. We have developed XCMS Online, a platform that enables online metabolomics data processing and interpretation. A systems biology workflow recently implemented within XCMS Online enables rapid metabolic pathway mapping using raw metabolomics data for investigating dysregulated metabolic processes. In addition, this platform supports integration of multi-omic (such as genomic and proteomic) data to garner further systems-wide mechanistic insight. Here, we provide an in-depth procedure showing how to effectively navigate and use the systems biology workflow within XCMS Online without a priori knowledge of the platform, including uploading liquid chromatography (LC)-mass spectrometry (MS) data from metabolite-extracted biological samples, defining the job parameters to identify features, correcting for retention time deviations, conducting statistical analysis of features between sample classes and performing predictive metabolic pathway analysis. Additional multi-omics data can be uploaded and overlaid with previously identified pathways to enhance systems-wide analysis of the observed dysregulations. We also describe unique visualization tools to assist in elucidation of statistically significant dysregulated metabolic pathways. Parameter input takes 5-10 min, depending on user experience; data processing typically takes 1-3 h, and data analysis takes â^¼30 min.
Journal Article
Benchmarking of analysis strategies for data-independent acquisition proteomics using a large-scale dataset comprising inter-patient heterogeneity
by
Fröhlich, Klemens
,
Fahrner, Matthias
,
Vogele, Daniel
in
631/114/1314
,
631/1647/296
,
631/45/475
2022
Numerous software tools exist for data-independent acquisition (DIA) analysis of clinical samples, necessitating their comprehensive benchmarking. We present a benchmark dataset comprising real-world inter-patient heterogeneity, which we use for in-depth benchmarking of DIA data analysis workflows for clinical settings. Combining spectral libraries, DIA software, sparsity reduction, normalization, and statistical tests results in 1428 distinct data analysis workflows, which we evaluate based on their ability to correctly identify differentially abundant proteins. From our dataset, we derive bootstrap datasets of varying sample sizes and use the whole range of bootstrap datasets to robustly evaluate each workflow. We find that all DIA software suites benefit from using a gas-phase fractionated spectral library, irrespective of the library refinement used. Gas-phase fractionation-based libraries perform best against two out of three reference protein lists. Among all investigated statistical tests non-parametric permutation-based statistical tests consistently perform best.
Data independent acquisition (DIA) has been gaining momentum in clinical proteomics. Here, the authors create a benchmark dataset comprising inter-patient heterogeneity to compare popular DIA data analysis workflows for identifying differentially abundant proteins.
Journal Article
Microbiome meta-analysis and cross-disease comparison enabled by the SIAMCAT machine learning toolbox
by
Karcher, Nicolai
,
Zeller, Georg
,
Zych, Konrad
in
Algorithms
,
Animal Genetics and Genomics
,
Bioinformatics
2021
The human microbiome is increasingly mined for diagnostic and therapeutic biomarkers using machine learning (ML). However, metagenomics-specific software is scarce, and overoptimistic evaluation and limited cross-study generalization are prevailing issues. To address these, we developed SIAMCAT, a versatile R toolbox for ML-based comparative metagenomics. We demonstrate its capabilities in a meta-analysis of fecal metagenomic studies (10,803 samples). When naively transferred across studies, ML models lost accuracy and disease specificity, which could however be resolved by a novel training set augmentation strategy. This reveals some biomarkers to be disease-specific, with others shared across multiple conditions. SIAMCAT is freely available from
siamcat.embl.de
.
Journal Article
MetaboAnalystR 4.0: a unified LC-MS workflow for global metabolomics
2024
The wide applications of liquid chromatography - mass spectrometry (LC-MS) in untargeted metabolomics demand an easy-to-use, comprehensive computational workflow to support efficient and reproducible data analysis. However, current tools were primarily developed to perform specific tasks in LC-MS based metabolomics data analysis. Here we introduce MetaboAnalystR 4.0 as a streamlined pipeline covering raw spectra processing, compound identification, statistical analysis, and functional interpretation. The key features of MetaboAnalystR 4.0 includes an auto-optimized feature detection and quantification algorithm for LC-MS1 spectra processing, efficient MS2 spectra deconvolution and compound identification for data-dependent or data-independent acquisition, and more accurate functional interpretation through integrated spectral annotation. Comprehensive validation studies using LC-MS1 and MS2 spectra obtained from standards mixtures, dilution series and clinical metabolomics samples have shown its excellent performance across a wide range of common tasks such as peak picking, spectral deconvolution, and compound identification with good computing efficiency. Together with its existing statistical analysis utilities, MetaboAnalystR 4.0 represents a significant step toward a unified, end-to-end workflow for LC-MS based global metabolomics in the open-source R environment.
Several bottlenecks exist in metabolomics data analysis. Here, the authors present MetaboAnalystR 4.0 as a unified workflow for LC-MS untargeted metabolomics. It highlights significant improvements in LC-MS2 spectral processing and functional analysis, providing an end-to-end computational pipeline.
Journal Article