Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
2,562
result(s) for
"Stability Selection"
Sort by:
A Stability-Oriented Biomarker Selection Framework Synergistically Driven by Robust Rank Aggregation and L1-Sparse Modeling
2025
Background: In high-dimensional, small-sample omics studies such as metabolomics, feature selection not only determines the discriminative performance of classification models but also directly affects the reproducibility and translational value of candidate biomarkers. However, most existing methods primarily optimize classification accuracy and treat stability as a post hoc diagnostic, leading to considerable fluctuations in selected feature sets under different data splits or mild perturbations. Methods: To address this issue, this study proposes FRL-TSFS, a feature selection framework synergistically driven by filter-based Robust Rank Aggregation and L1-sparse modeling. Five complementary filter methods—variance thresholding, chi-square test, mutual information, ANOVA F test, and ReliefF—are first applied in parallel to score features, and Robust Rank Aggregation (RRA) is then used to obtain a consensus feature ranking that is less sensitive to the bias of any single scoring criterion. An L1-regularized logistic regression model is subsequently constructed on the candidate feature subset defined by the RRA ranking to achieve task-coupled sparse selection, thereby linking feature selection stability, feature compression, and classification performance. Results: FRL-TSFS was evaluated on six representative metabolomics and gene expression datasets under a mildly perturbed scenario induced by 10-fold cross-validation, and its performance was compared with multiple baselines using the Extended Kuncheva Index (EKI), Accuracy, and F1-score. The results show that RRA substantially improves ranking stability compared with conventional aggregation strategies without degrading classification performance, while the full FRL-TSFS framework consistently attains higher EKI values than the other feature selection schemes, markedly reduces the number of selected features to several tens of metabolites or genes, and maintains competitive classification performance. Conclusions: These findings indicate that FRL-TSFS can generate compact, reproducible, and interpretable biomarker panels, providing a practical analysis framework for stability-oriented feature selection and biomarker discovery in untargeted metabolomics.
Journal Article
Genome‐wide association study of café‐au‐lait macule number in neurofibromatosis type 1
by
Bass, Sara
,
Wilson, Alexander F.
,
Hyland, Paula L.
in
Adult
,
Association analysis
,
Cafe-au-Lait Spots - genetics
2020
Background Neurofibromatosis type 1 (NF1) is a tumor‐predisposition disorder that arises due to pathogenic variants in tumor suppressor NF1. NF1 has variable expressivity that may be due, at least in part, from heritable elements such as modifier genes; however, few genetic modifiers have been identified to date. Methods In this study, we performed a genome‐wide association analysis of the number of café‐au‐lait macules (CALM) that are considered a tumor‐like trait as a clinical phenotype modifying NF1. Results A borderline genome‐wide significant association was identified in the discovery cohort (CALM1, N = 112) between CALM number and rs12190451 (and rs3799603, r2 = 1.0; p = 7.4 × 10−8) in the intronic region of RPS6KA2. Although, this association was not replicated in the second cohort (CALM2, N = 59) and a meta‐analysis did not show significantly associated variants in this region, a significant corroboration score (0.72) was obtained for the RPS6KA2 signal in the discovery cohort (CALM1) using Complementary Pairs Stability Selection for Genome‐Wide Association Studies (ComPaSS‐GWAS) analysis, suggesting that the lack of replication may be due to heterogeneity of the cohorts rather than type I error. Conclusion rs12190451 is located in a melanocyte‐specific enhancer and may influence RPS6KA2 expression in melanocytes—warranting further functional studies. In this study, we performed a genome‐wide association analysis of the number of café‐au‐lait macules (CALM) that are considered a tumor‐like trait as a clinical phenotype modifying neurofibromatosis type 1. A borderline genome‐wide significant association was identified in the discovery cohort (CALM1, N = 112) between CALM number and rs12190451 (p = 7.4 × 10–8) in the intronic region of RPS6KA2. rs12190451 is located in a melanocyte‐specific enhancer and may influence expression of RPS6KA2, a biologically compelling candidate as a NF1 genetic modifier, since the protein is phosphorylated and activated by RAS‐MAPK pathway kinases ERK1/2 that act downstream of RAS and neurofibromin.
Journal Article
Variable selection with error control: another look at stability selection
by
Shah, Rajen D.
,
Samworth, Richard J.
in
Algorithms
,
Complementary pairs stability selection
,
Constrictions
2013
Stability selection was recently introduced by Meinshausen and Bühlmann as a very general technique designed to improve the performance of a variable selection algorithm. It is based on aggregating the results of applying a selection procedure to subsamples of the data. We introduce a variant, called complementary pairs stability selection, and derive bounds both on the expected number of variables included by complementary pairs stability selection that have low selection probability under the original procedure, and on the expected number of high selection probability variables that are excluded. These results require no (e.g. exchangeability) assumptions on the underlying model or on the quality of the original selection procedure. Under reasonable shape restrictions, the bounds can be further tightened, yielding improved error control, and therefore increasing the applicability of the methodology.
Journal Article
Stability selection
2010
Estimation of structure, such as in variable selection, graphical modelling or cluster analysis, is notoriously difficult, especially for high dimensional data. We introduce stability selection. It is based on subsampling in combination with (high dimensional) selection algorithms. As such, the method is extremely general and has a very wide range of applicability. Stability selection provides finite sample control for some error rates of false discoveries and hence a transparent principle to choose a proper amount of regularization for structure estimation. Variable selection and structure estimation improve markedly for a range of selection methods if stability selection is applied. We prove for the randomized lasso that stability selection will be variable selection consistent even if the necessary conditions for consistency of the original lasso method are violated. We demonstrate stability selection for variable selection and Gaussian graphical modelling, using real and simulated data.
Journal Article
Stability selection enables robust learning of differential equations from limited noisy data
by
Müller, Christian L.
,
Cheeseman, Bevan L.
,
Maddu, Suryanarayana
in
Differential Equations
,
Machine Learning
,
Par proteins
2022
We present a statistical learning framework for robust identification of differential equations from noisy spatio-temporal data. We address two issues that have so far limited the application of such methods, namely their robustness against noise and the need for manual parameter tuning, by proposing stability-based model selection to determine the level of regularization required for reproducible inference. This avoids manual parameter tuning and improves robustness against noise in the data. Our stability selection approach, termed PDE-STRIDE, can be combined with any sparsity-promoting regression method and provides an interpretable criterion for model component importance. We show that the particular combination of stability selection with the iterative hard-thresholding algorithm from compressed sensing provides a fast and robust framework for equation inference that outperforms previous approaches with respect to accuracy, amount of data required, and robustness. We illustrate the performance of PDE-STRIDE on a range of simulated benchmark problems, and we demonstrate the applicability of PDE-STRIDE on real-world data by considering purely data-driven inference of the protein interaction network for embryonic polarization in Caenorhabditis elegans. Using fluorescence microscopy images of C. elegans zygotes as input data, PDE-STRIDE is able to learn the molecular interactions of the proteins.
Journal Article
Feature Stability as a Trust Layer for Feature Selection: Resampling-Based Recurrence Profiles Beyond Predictive Performance
by
Elmakias, Itamar
,
Kolsky, Dor
,
Vilenchik, Dan
in
chance-corrected stability
,
Dimensional analysis
,
Feature selection
2026
Feature selection in high-dimensional studies is conventionally evaluated by the predictive performance of the features it returns, but predictive performance does not indicate whether the same subset would be selected again under a reasonable perturbation of the data. We propose a feature-stability profile, reported as a diagnostic layer beside predictive performance rather than in place of it. The profile is assembled from established quantities: per-feature selection frequency across repeated stratified resamples, a chance-corrected stability summary, recurrent sets reported across a sweep of descriptive cutoffs, a random-selection baseline, and the held-out predictive performance recorded on the same resamples. We examine it in two settings. In a controlled synthetic study, the informative support is planted by construction, so recovery can be measured directly; in an illustrative application across high-dimensional binary datasets, no such support exists, and a broader exploratory roster is reported as supporting results. Under planted support, predictive performance, subset stability, and support recovery can diverge rather than decline together: at an intermediate signal level, performance can remain relatively preserved while exact recovery falls, and some low exact overlap reflects substitution among redundant alternatives. On real data, recurrent features are treated as recurrent candidates, not recovered or validated features; selectors reaching near-equal area under the receiver operating characteristic curve (AUC) can differ about twofold in chance-corrected recurrence. The contribution is diagnostic and integrative, not a new selector, a new metric, a benchmark ranking, or an error-controlled procedure, making the reliability of a selected feature set visible rather than assumed.
Journal Article
Trimming stability selection increases variable selection robustness
2023
Contamination can severely distort an estimator unless the estimation procedure is suitably robust. This is a well-known issue and has been addressed in Robust Statistics, however, the relation of contamination and distorted variable selection has been rarely considered in the literature. As for variable selection, many methods for sparse model selection have been proposed, including the Stability Selection which is a meta-algorithm based on some variable selection algorithm in order to immunize against particular data configurations. We introduce the variable selection breakdown point that quantifies the number of cases resp. cells that have to be contaminated in order to let no relevant variable be detected. We show that particular outlier configurations can completely mislead model selection. We combine the variable selection breakdown point with resampling, resulting in the Stability Selection breakdown point that quantifies the robustness of Stability Selection. We propose a trimmed Stability Selection which only aggregates the models with the best performance so that, heuristically, models computed on heavily contaminated resamples should be trimmed away. An extensive simulation study with non-robust regression and classification algorithms as well as with two robust regression algorithms reveals both the potential of our approach to boost the model selection robustness as well as the fragility of variable selection using non-robust algorithms, even for an extremely small cell-wise contamination rate.
Journal Article
Stability-driven nonnegative matrix factorization to interpret spatial gene expression and build local gene networks
by
Celniker, Susan E.
,
Frise, Erwin
,
Yu, Bin
in
Animals
,
BASIC BIOLOGICAL SCIENCES
,
Biological Sciences
2016
Spatial gene expression patterns enable the detection of local covariability and are extremely useful for identifying local gene interactions during normal development. The abundance of spatial expression data in recent years has led to the modeling and analysis of regulatory networks. The inherent complexity of such data makes it a challenge to extract biological information. We developed staNMF, a method that combines a scalable implementation of nonnegative matrix factorization (NMF) with a new stability-driven model selection criterion. When applied to a set of Drosophila early embryonic spatial gene expression images, one of the largest datasets of its kind, staNMF identified 21 principal patterns (PP). Providing a compact yet biologically interpretable representation of Drosophila expression patterns, PP are comparable to a fate map generated experimentally by laser ablation and show exceptional promise as a data-driven alternative to manual annotations. Our analysis mapped genes to cell-fate programs and assigned putative biological roles to uncharacterized genes. Finally, we used the PP to generate local transcription factor regulatory networks. Spatially local correlation networks were constructed for six PP that span along the embryonic anterior–posterior axis. Using a two-tail 5% cutoff on correlation, we reproduced 10 of the 11 links in the well-studied gap gene network. The performance of PP with the Drosophila data suggests that staNMF provides informative decompositions and constitutes a useful computational lens through which to extract biological insight from complex and often noisy gene expression data.
Journal Article
Landscape predictors of pathogen prevalence and range contractions in US bumblebees
by
Adler, Lynn S.
,
Irwin, Rebecca E.
,
Urbanowicz, Christine
in
Agriculture - methods
,
Animal Distribution
,
Animals
2017
Several species of bumblebees have recently experienced range contractions and possible extinctions. While threats to bees are numerous, few analyses have attempted to understand the relative importance of multiple stressors. Such analyses are critical for prioritizing conservation strategies. Here, we describe a landscape analysis of factors predicted to cause bumblebee declines in the USA. We quantified 24 habitat, land-use and pesticide usage variables across 284 sampling locations, assessing which variables predicted pathogen prevalence and range contractions via machine learning model selection techniques. We found that greater usage of the fungicide chlorothalonil was the best predictor of pathogen (Nosema bombi) prevalence in four declining species of bumblebees. Nosema bombi has previously been found in greater prevalence in some declining US bumblebee species compared to stable species. Greater usage of total fungicides was the strongest predictor of range contractions in declining species, with bumblebees in the northern USA experiencing greater likelihood of loss from previously occupied areas. These results extend several recent laboratory and semi-field studies that have found surprising links between fungicide exposure and bee health. Specifically, our data suggest landscape-scale connections between fungicide usage, pathogen prevalence and declines of threatened and endangered bumblebees.
Journal Article
Controlling false discoveries in high-dimensional situations: boosting with stability selection
2015
Background
Modern biotechnologies often result in high-dimensional data sets with many more variables than observations (
n
≪
p
). These data sets pose new challenges to statistical analysis: Variable selection becomes one of the most important tasks in this setting. Similar challenges arise if in modern data sets from observational studies, e.g., in ecology, where flexible, non-linear models are fitted to high-dimensional data. We assess the recently proposed flexible framework for variable selection called stability selection. By the use of resampling procedures, stability selection adds a finite sample error control to high-dimensional variable selection procedures such as Lasso or boosting. We consider the combination of boosting and stability selection and present results from a detailed simulation study that provide insights into the usefulness of this combination. The interpretation of the used error bounds is elaborated and insights for practical data analysis are given.
Results
Stability selection with boosting was able to detect influential predictors in high-dimensional settings while controlling the given error bound in various simulation scenarios. The dependence on various parameters such as the sample size, the number of truly influential variables or tuning parameters of the algorithm was investigated. The results were applied to investigate phenotype measurements in patients with autism spectrum disorders using a log-linear interaction model which was fitted by boosting. Stability selection identified five differentially expressed amino acid pathways.
Conclusion
Stability selection is implemented in the freely available
R
package stabs (
http://CRAN.R-project.org/package=stabs
). It proved to work well in high-dimensional settings with more predictors than observations for both, linear and additive models. The original version of stability selection, which controls the per-family error rate, is quite conservative, though, this is much less the case for its improvement, complementary pairs stability selection. Nevertheless, care should be taken to appropriately specify the error bound.
Journal Article