Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
      More Filters
      Clear All
      More Filters
      Source
    • Language
3,837 result(s) for "biomarker selection"
Sort by:
Identification of biomarkers predictive of metastasis development in early-stage colorectal cancer using network-based regularization
Colorectal cancer (CRC) is the third most common cancer and the second most deathly worldwide. It is a very heterogeneous disease that can develop via distinct pathways where metastasis is the primary cause of death. Therefore, it is crucial to understand the molecular mechanisms underlying metastasis. RNA-sequencing is an essential tool used for studying the transcriptional landscape. However, the high-dimensionality of gene expression data makes selecting novel metastatic biomarkers problematic. To distinguish early-stage CRC patients at risk of developing metastasis from those that are not, three types of binary classification approaches were used: (1) classification methods (decision trees, linear and radial kernel support vector machines, logistic regression, and random forest) using differentially expressed genes (DEGs) as input features; (2) regularized logistic regression based on the Elastic Net penalty and the proposed iTwiner—a network-based regularizer accounting for gene correlation information; and (3) classification methods based on the genes pre-selected using regularized logistic regression. Classifiers using the DEGs as features showed similar results, with random forest showing the highest accuracy. Using regularized logistic regression on the full dataset yielded no improvement in the methods’ accuracy. Further classification using the pre-selected genes found by different penalty factors, instead of the DEGs, significantly improved the accuracy of the binary classifiers. Moreover, the use of network-based correlation information (iTwiner) for gene selection produced the best classification results and the identification of more stable and robust gene sets. Some are known to be tumor suppressor genes ( OPCML-IT2 ), to be related to resistance to cancer therapies ( RAC1P3 ), or to be involved in several cancer processes such as genome stability ( XRCC6P2 ), tumor growth and metastasis ( MIR602 ) and regulation of gene transcription ( NME2P2 ). We show that the classification of CRC patients based on pre-selected features by regularized logistic regression is a valuable alternative to using DEGs, significantly increasing the models’ predictive performance. Moreover, the use of correlation-based penalization for biomarker selection stands as a promising strategy for predicting patients’ groups based on RNA-seq data.
SHAP-based binarization enhances metataxonomic machine learning with application to gut microbiota of inflammatory bowel disease
Machine learning has been increasingly applied to microbiome data for biomarker discovery. However, microbiome datasets are typically high-dimensional, sparse, and correlated, which makes model training challenging and prone to overfitting. Previous studies have also reported that microbiome features exhibit binary-like characteristics, and that binarization does not necessarily reduce predictive performance. This observation motivated our work. Building on this idea, we propose a SHAP-based binarization pipeline. We first trained several machine learning models on raw continuous data and selected the best-performing model (random forest). Using SHAP values derived from the training set, we determined feature-specific thresholds that best separated positive and negative contributions. The dataset was then binarized using these thresholds and new models were trained on the transformed data. We evaluated this approach on gut microbiome abundance data (283 species, 220 genera, 1,569 individuals) to classify inflammatory bowel disease (IBD) versus healthy controls. The SHAP-based binarization consistently improved classification performance and interpretability compared with both continuous data and zero-threshold binarization. The best model’s Matthews correlation coefficient increased from 0.884 to 0.928, with the largest improvements observed in non-tree-based models such as logistic regression and neural networks. SHAP summary plots also revealed clearer feature patterns, and biomarker rankings were more stable. In addition, the pipeline enabled us to identify a concise set of 17 microbial biomarkers associated with IBD. This study introduces a novel approach for microbiome data analysis by explicitly linking binarization thresholds to SHAP-derived feature contributions. Our approach was grounded in the observation of binary-like patterns revealed through SHAP values. Furthermore, although binarization inevitably raises concerns about information loss, our evaluation confirmed improvements not only in predictive performance but also in interpretability and biomarker stability, providing a broader validation of robustness. These findings highlight SHAP-based binarization as an effective strategy for high-dimensional microbiome data, with broad applicability and opportunities for future extension.
Tumor‐Specific Success Probabilities and Factors Associated With Phase III Trials in Oncology Drug Development
Despite advances in novel therapeutic modalities, such as molecularly targeted agents and immunotherapies, the probability of success in Phase III oncology trials remains low. This study quantitatively characterized success probabilities of Phase III trials in oncology drug development and evaluated the factors associated with trial success. Phase III interventional oncology drug trials registered at ClinicalTrials.gov between 2007 and 2023 with publicly available primary endpoint results were included. Trial success was defined as the achievement of at least one primary endpoint, and tumor‐specific success probabilities were calculated. Multivariable logistic regression analyses were conducted to assess the association between trial success and key trial characteristics. We analyzed 824 trials (358 successful and 466 unsuccessful). Overall success probabilities were comparable between solid tumors and hematologic malignancies, although substantial heterogeneity was observed across solid tumors, with particularly low success probabilities for central nervous system tumors and pancreatic cancer. Multivariable analyses showed that biomarker‐based patient selection, more recently initiated trials, line of therapy, and the number of primary endpoints were associated with trial success. Trials evaluating molecularly targeted therapies and those with short‐term evaluable endpoints showed higher success probabilities, whereas trials evaluating chemotherapy or assessing overall or event‐free survival endpoints showed lower success probabilities. Phase III trial success in oncology is associated with tumor‐specific characteristics and development‐stage factors, including biomarker‐based patient selection and trial design. These results provide quantitative evidence to inform decision‐making and trial design in oncology drug development. Study Highlights What is the current knowledge on the topic? ○Phase III oncology trials have a relatively low probability of success, and trial failure has been attributed to biological complexity, tumor heterogeneity, and challenges in late‐stage clinical development. However, relatively few studies have quantitatively characterized tumor‐specific success probabilities and comprehensively evaluated trial design‐related factors across oncological indications. What question did this study address? ○This study examined tumor‐specific success probabilities and quantitatively evaluated the factors associated with Phase III trial success in oncology drug development using publicly available trial data and multivariable logistic regression analysis. What does this study add to our knowledge? ○This study demonstrates substantial heterogeneity in Phase III trial success probabilities across tumor types and identifies key developmental stage factors associated with trial success, including biomarker‐based patient selection, therapeutic modality, line of therapy, and primary endpoint selection. How might this change clinical pharmacology or translational science? ○These findings provide quantitative evidence to inform decision‐making and trial design in late‐stage oncology drug development, supporting more rational and efficient development strategies for precision oncology.
Integrating Multi-Omics Analysis for Enhanced Diagnosis and Treatment of Glioblastoma: A Comprehensive Data-Driven Approach
The most aggressive primary malignant brain tumor in adults is glioblastoma (GBM), which has poor overall survival (OS). There is a high relapse rate among patients with GBM despite maximally safe surgery, radiation therapy, temozolomide (TMZ), and aggressive treatment. Hence, there is an urgent and unmet clinical need for new approaches to managing GBM. The current study identified modules (MYC, EGFR, PIK3CA, SUZ12, and SPRK2) involved in GBM disease through the NeDRex plugin. Furthermore, hub genes were identified in a comprehensive interaction network containing 7560 proteins related to GBM disease and 3860 proteins associated with signaling pathways involved in GBM. By integrating the results of the analyses mentioned above and again performing centrality analysis, eleven key genes involved in GBM disease were identified. ProteomicsDB and Gliovis databases were used for determining the gene expression in normal and tumor brain tissue. The NetworkAnalyst and the mGWAS-Explorer tools identified miRNAs, SNPs, and metabolites associated with these 11 genes. Moreover, a literature review of recent studies revealed other lists of metabolites related to GBM disease. The enrichment analysis of identified genes, miRNAs, and metabolites associated with GBM disease was performed using ExpressAnalyst, miEAA, and MetaboAnalyst tools. Further investigation of metabolite roles in GBM was performed using pathway, joint pathway, and network analyses. The results of this study allowed us to identify 11 genes (UBC, HDAC1, CTNNB1, TRIM28, CSNK2A1, RBBP4, TP53, APP, DAB1, PINK1, and RELN), five miRNAs (hsa-mir-221-3p, hsa-mir-30a-5p, hsa-mir-15a-5p, hsa-mir-130a-3p, and hsa-let-7b-5p), six metabolites (HDL, N6-acetyl-L-lysine, cholesterol, formate, N, N-dimethylglycine/xylose, and X2. piperidinone) and 15 distinct signaling pathways that play an indispensable role in GBM disease development. The identified top genes, miRNAs, and metabolite signatures can be targeted to establish early diagnostic methods and plan personalized GBM treatment strategies.
A Stability-Oriented Biomarker Selection Framework Synergistically Driven by Robust Rank Aggregation and L1-Sparse Modeling
Background: In high-dimensional, small-sample omics studies such as metabolomics, feature selection not only determines the discriminative performance of classification models but also directly affects the reproducibility and translational value of candidate biomarkers. However, most existing methods primarily optimize classification accuracy and treat stability as a post hoc diagnostic, leading to considerable fluctuations in selected feature sets under different data splits or mild perturbations. Methods: To address this issue, this study proposes FRL-TSFS, a feature selection framework synergistically driven by filter-based Robust Rank Aggregation and L1-sparse modeling. Five complementary filter methods—variance thresholding, chi-square test, mutual information, ANOVA F test, and ReliefF—are first applied in parallel to score features, and Robust Rank Aggregation (RRA) is then used to obtain a consensus feature ranking that is less sensitive to the bias of any single scoring criterion. An L1-regularized logistic regression model is subsequently constructed on the candidate feature subset defined by the RRA ranking to achieve task-coupled sparse selection, thereby linking feature selection stability, feature compression, and classification performance. Results: FRL-TSFS was evaluated on six representative metabolomics and gene expression datasets under a mildly perturbed scenario induced by 10-fold cross-validation, and its performance was compared with multiple baselines using the Extended Kuncheva Index (EKI), Accuracy, and F1-score. The results show that RRA substantially improves ranking stability compared with conventional aggregation strategies without degrading classification performance, while the full FRL-TSFS framework consistently attains higher EKI values than the other feature selection schemes, markedly reduces the number of selected features to several tens of metabolites or genes, and maintains competitive classification performance. Conclusions: These findings indicate that FRL-TSFS can generate compact, reproducible, and interpretable biomarker panels, providing a practical analysis framework for stability-oriented feature selection and biomarker discovery in untargeted metabolomics.
Favoring the hierarchical constraint in penalized survival models for randomized trials in precision medicine
Background The research of biomarker-treatment interactions is commonly investigated in randomized clinical trials (RCT) for improving medicine precision. The hierarchical interaction constraint states that an interaction should only be in a model if its main effects are also in the model. However, this constraint is not guaranteed in the standard penalized statistical approaches. We aimed to find a compromise for high-dimensional data between the need for sparse model selection and the need for the hierarchical constraint. Results To favor the property of the hierarchical interaction constraint, we proposed to create groups composed of the biomarker main effect and its interaction with treatment and to perform the bi-level selection on these groups. We proposed two weighting approaches (Single Wald (SW) and likelihood ratio test (LRT)) for the adaptive lasso method. The selection performance of these two approaches is compared to alternative lasso extensions (adaptive lasso with ridge-based weights, composite Minimax Concave Penalty, group exponential lasso and Sparse Group Lasso) through a simulation study. A RCT (NSABP B-31) randomizing 1574 patients (431 events) with early breast cancer aiming to evaluate the effect of adjuvant trastuzumab on distant-recurrence free survival with expression data from 462 genes measured in the tumour will serve for illustration. The simulation study illustrates that the adaptive lasso LRT and SW, and the group exponential lasso favored the hierarchical interaction constraint. Overall, in the alternative scenarios, they had the best balance of false discovery and false negative rates for the main effects of the selected interactions. For NSABP B-31, 12 gene-treatment interactions were identified more than 20% by the different methods. Among them, the adaptive lasso (SW) approach offered the best trade-off between a high number of selected gene-treatment interactions and a high proportion of selection of both the gene-treatment interaction and its main effect. Conclusions Adaptive lasso with Single Wald and likelihood ratio test weighting and the group exponential lasso approaches outperformed their competitors in favoring the hierarchical constraint of the biomarker-treatment interaction. However, the performance of the methods tends to decrease in the presence of prognostic biomarkers.
Toxicity Evaluation and Biomarker Selection with Validated Reference Gene in Embryonic Zebrafish Exposed to Mitoxantrone
Notwithstanding the widespread use and promising clinical value of chemotherapy, the pharmacokinetics, toxicology, and mechanism of mitoxantrone remains unclear. To promote the clinical value in the treatment of human diseases and the exploration of potential subtle effects of mitoxantrone, zebrafish embryos were employed to evaluate toxicity with validated reference genes based on independent stability evaluation programs. The most stable and recommended reference gene was gapdh, followed by tubα1b, for the 48 h post fertilization (hpf) zebrafish embryo mitoxantrone test, while both eef1a1l1 and rpl13α were recommended as reference genes for the 96 hpf zebrafish embryo mitoxantrone test. With gapdh as an internal control, we analyzed the mRNA levels of representative hepatotoxicity biomarkers, including fabp10a, gclc, gsr, nqo1, cardiotoxicity biomarker erg, and neurotoxicity biomarker gfap in the 48 hpf embryo mitoxantrone test. The mRNA levels of gclc, gsr, and gfap increased significantly in 10 and 50 μg/L mitoxantrone-treated 48 hpf embryos, while the transcript levels of fabp10a decreased in a dose-dependent manner, indicating that mitoxantrone induced hepatotoxicity and neurotoxicity. Liver hematoxylin–eosin staining and the spontaneous movement of embryos confirmed the results. Thus, the present research suggests that mitoxantrone induces toxicity during the development of the liver and nervous system in zebrafish embryos and that fabp10a is recommended as a potential biomarker for hepatotoxicity in zebrafish embryos. Additionally, gapdh is proposed as a reference gene for the 48 hpf zebrafish embryo mitoxantrone toxicity test, while eef1a1l1 and rpl13α are proposed as that for the 96 hpf test.
Accounting for grouped predictor variables or pathways in high-dimensional penalized Cox regression models
Background The standard lasso penalty and its extensions are commonly used to develop a regularized regression model while selecting candidate predictor variables on a time-to-event outcome in high-dimensional data. However, these selection methods focus on a homogeneous set of variables and do not take into account the case of predictors belonging to functional groups; typically, genomic data can be grouped according to biological pathways or to different types of collected data. Another challenge is that the standard lasso penalisation is known to have a high false discovery rate. Results We evaluated different penalizations in a Cox model to select grouped variables in order to further penalize variables that, in addition to having a low effect, belong to a group with a low overall effect; and to favor the selection of variables that, in addition to having a large effect, belong to a group with a large overall effect. We considered the case of prespecified and disjoint groups and proposed diverse weights for the adaptive lasso method. In particular we proposed the product Max Single Wald by Single Wald weighting (MSW*SW) which takes into account the information of the group to which it belongs and of this biomarker. Through simulations, we compared the selection and prediction ability of our approach with the standard lasso, the composite Minimax Concave Penalty (cMCP), the group exponential lasso (gel), the Integrative L 1-Penalized Regression with Penalty Factors (IPF-Lasso), and the Sparse Group Lasso (SGL) methods. In addition, we illustrated the methods using gene expression data of 614 breast cancer patients. Conclusions The adaptive lasso with the MSW*SW weighting method incorporates both the information in the grouping structure and the individual variable. It outperformed the competitors by reducing the false discovery rate without severely increasing the false negative rate.
LANDMark: an ensemble approach to the supervised selection of biomarkers in high-throughput sequencing data
Background Identification of biomarkers, which are measurable characteristics of biological datasets, can be challenging. Although amplicon sequence variants (ASVs) can be considered potential biomarkers, identifying important ASVs in high-throughput sequencing datasets is challenging. Noise, algorithmic failures to account for specific distributional properties, and feature interactions can complicate the discovery of ASV biomarkers. In addition, these issues can impact the replicability of various models and elevate false-discovery rates. Contemporary machine learning approaches can be leveraged to address these issues. Ensembles of decision trees are particularly effective at classifying the types of data commonly generated in high-throughput sequencing (HTS) studies due to their robustness when the number of features in the training data is orders of magnitude larger than the number of samples. In addition, when combined with appropriate model introspection algorithms, machine learning algorithms can also be used to discover and select potential biomarkers. However, the construction of these models could introduce various biases which potentially obfuscate feature discovery. Results We developed a decision tree ensemble, LANDMark, which uses oblique and non-linear cuts at each node. In synthetic and toy tests LANDMark consistently ranked as the best classifier and often outperformed the Random Forest classifier. When trained on the full metabarcoding dataset obtained from Canada’s Wood Buffalo National Park, LANDMark was able to create highly predictive models and achieved an overall balanced accuracy score of 0.96 ± 0.06. The use of recursive feature elimination did not impact LANDMark’s generalization performance and, when trained on data from the BE amplicon, it was able to outperform the Linear Support Vector Machine, Logistic Regression models, and Stochastic Gradient Descent models ( p  ≤ 0.05). Finally, LANDMark distinguishes itself due to its ability to learn smoother non-linear decision boundaries. Conclusions Our work introduces LANDMark, a meta-classifier which blends the characteristics of several machine learning models into a decision tree and ensemble learning framework. To our knowledge, this is the first study to apply this type of ensemble approach to amplicon sequencing data and we have shown that analyzing these datasets using LANDMark can produce highly predictive and consistent models.
The role of network science in glioblastoma
Network science has long been recognized as a well-established discipline across many biological domains. In the particular case of cancer genomics, network discovery is challenged by the multitude of available high-dimensional heterogeneous views of data. Glioblastoma (GBM) is an example of such a complex and heterogeneous disease that can be tackled by network science. Identifying the architecture of molecular GBM networks is essential to understanding the information flow and better informing drug development and pre-clinical studies. Here, we review network-based strategies that have been used in the study of GBM, along with the available software implementations for reproducibility and further testing on newly coming datasets. Promising results have been obtained from both bulk and single-cell GBM data, placing network discovery at the forefront of developing a molecularly-informed-based personalized medicine.