Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
      More Filters
      Clear All
      More Filters
      Source
    • Language
3,877 result(s) for "Selection methods (Regression analysis)"
Sort by:
Evaluating variable selection methods for multivariable regression models: A simulation study protocol
Researchers often perform data-driven variable selection when modeling the associations between an outcome and multiple independent variables in regression analysis. Variable selection may improve the interpretability, parsimony and/or predictive accuracy of a model. Yet variable selection can also have negative consequences, such as false exclusion of important variables or inclusion of noise variables, biased estimation of regression coefficients, underestimated standard errors and invalid confidence intervals, as well as model instability. While the potential advantages and disadvantages of variable selection have been discussed in the literature for decades, few large-scale simulation studies have neutrally compared data-driven variable selection methods with respect to their consequences for the resulting models. We present the protocol for a simulation study that will evaluate different variable selection methods: forward selection, stepwise forward selection, backward elimination, augmented backward elimination, univariable selection, univariable selection followed by backward elimination, and penalized likelihood approaches (Lasso, relaxed Lasso, adaptive Lasso). These methods will be compared with respect to false inclusion and/or exclusion of variables, consequences on bias and variance of the estimated regression coefficients, the validity of the confidence intervals for the coefficients, the accuracy of the estimated variable importance ranking, and the predictive performance of the selected models. We consider both linear and logistic regression in a low-dimensional setting (20 independent variables with 10 true predictors and 10 noise variables). The simulation will be based on real-world data from the National Health and Nutrition Examination Survey (NHANES). Publishing this study protocol ahead of performing the simulation increases transparency and allows integrating the perspective of other experts into the study design.
Genomic selection using random regressions on known and latent environmental covariates
Key messageThe integration of known and latent environmental covariates within a single-stage genomic selection approach provides breeders with an informative and practical framework to utilise genotype by environment interaction for prediction into current and future environments.This paper develops a single-stage genomic selection approach which integrates known and latent environmental covariates within a special factor analytic framework. The factor analytic linear mixed model of Smith et al. (2001) is an effective method for analysing multi-environment trial (MET) datasets, but has limited practicality since the underlying factors are latent so the modelled genotype by environment interaction (GEI) is observable, rather than predictable. The advantage of using random regressions on known environmental covariates, such as soil moisture and daily temperature, is that the modelled GEI becomes predictable. The integrated factor analytic linear mixed model (IFA-LMM) developed in this paper includes a model for predictable and observable GEI in terms of a joint set of known and latent environmental covariates. The IFA-LMM is demonstrated on a late-stage cotton breeding MET dataset from Bayer CropScience. The results show that the known covariates predominately capture crossover GEI and explain 34.4% of the overall genetic variance. The most notable covariates are maximum downward solar radiation (10.1%), average cloud cover (4.5%) and maximum temperature (4.0%). The latent covariates predominately capture non-crossover GEI and explain 40.5% of the overall genetic variance. The results also show that the average prediction accuracy of the IFA-LMM is 0.02-0.10 higher than conventional random regression models for current environments and 0.06-0.24 higher for future environments. The IFA-LMM is therefore an effective method for analysing MET datasets which also utilises crossover and non-crossover GEI for genomic prediction into current and future environments. This is becoming increasingly important with the emergence of rapidly changing environments and climate change.
HDSI: High dimensional selection with interactions algorithm on feature selection and testing
Feature selection on high dimensional data along with the interaction effects is a critical challenge for classical statistical learning techniques. Existing feature selection algorithms such as random LASSO leverages LASSO capability to handle high dimensional data. However, the technique has two main limitations, namely the inability to consider interaction terms and the lack of a statistical test for determining the significance of selected features. This study proposes a High Dimensional Selection with Interactions (HDSI) algorithm, a new feature selection method, which can handle high-dimensional data, incorporate interaction terms, provide the statistical inferences of selected features and leverage the capability of existing classical statistical techniques. The method allows the application of any statistical technique like LASSO and subset selection on multiple bootstrapped samples; each contains randomly selected features. Each bootstrap data incorporates interaction terms for the randomly sampled features. The selected features from each model are pooled and their statistical significance is determined. The selected statistically significant features are used as the final output of the approach, whose final coefficients are estimated using appropriate statistical techniques. The performance of HDSI is evaluated using both simulated data and real studies. In general, HDSI outperforms the commonly used algorithms such as LASSO, subset selection, adaptive LASSO, random LASSO and group LASSO.
Randomized boosting with multivariable base-learners for high-dimensional variable selection and prediction
Background Statistical boosting is a computational approach to select and estimate interpretable prediction models for high-dimensional biomedical data, leading to implicit regularization and variable selection when combined with early stopping. Traditionally, the set of base-learners is fixed for all iterations and consists of simple regression learners including only one predictor variable at a time. Furthermore, the number of iterations is typically tuned by optimizing the predictive performance, leading to models which often include unnecessarily large numbers of noise variables. Results We propose three consecutive extensions of classical component-wise gradient boosting. In the first extension, called Subspace Boosting (SubBoost), base-learners can consist of several variables, allowing for multivariable updates in a single iteration. To compensate for the larger flexibility, the ultimate selection of base-learners is based on information criteria leading to an automatic stopping of the algorithm. As the second extension, Random Subspace Boosting (RSubBoost) additionally includes a random preselection of base-learners in each iteration, enabling the scalability to high-dimensional data. In a third extension, called Adaptive Subspace Boosting (AdaSubBoost), an adaptive random preselection of base-learners is considered, focusing on base-learners which have proven to be predictive in previous iterations. Simulation results show that the multivariable updates in the three subspace algorithms are particularly beneficial in cases of high correlations among signal covariates. In several biomedical applications the proposed algorithms tend to yield sparser models than classical statistical boosting, while showing a very competitive predictive performance also compared to penalized regression approaches like the (relaxed) lasso and the elastic net. Conclusions The proposed randomized boosting approaches with multivariable base-learners are promising extensions of statistical boosting, particularly suited for highly-correlated and sparse high-dimensional settings. The incorporated selection of base-learners via information criteria induces automatic stopping of the algorithms, promoting sparser and more interpretable prediction models.
Statistical model building: Background “knowledge” based on inappropriate preselection causes misspecification
Background Statistical model building requires selection of variables for a model depending on the model’s aim. In descriptive and explanatory models, a common recommendation often met in the literature is to include all variables in the model which are assumed or known to be associated with the outcome independent of their identification with data driven selection procedures. An open question is, how reliable this assumed “background knowledge” truly is. In fact, “known” predictors might be findings from preceding studies which may also have employed inappropriate model building strategies. Methods We conducted a simulation study assessing the influence of treating variables as “known predictors” in model building when in fact this knowledge resulting from preceding studies might be insufficient. Within randomly generated preceding study data sets, model building with variable selection was conducted. A variable was subsequently considered as a “known” predictor if a predefined number of preceding studies identified it as relevant. Results Even if several preceding studies identified a variable as a “true” predictor, this classification is often false positive. Moreover, variables not identified might still be truly predictive. This especially holds true if the preceding studies employed inappropriate selection methods such as univariable selection. Conclusions The source of “background knowledge” should be evaluated with care. Knowledge generated on preceding studies can cause misspecification.
Quantitative evaluation of recruitment strategies in a cluster-randomized trial: segmented regression and cost analysis from the AOK-Family + study
Background Recruitment of participants for preventive health intervention studies remains a significant challenge: Approximately 19% of studies are discontinued due to insufficient participant numbers, and one in three extends its recruitment period. This study aimed to examine the impact of recruiting measures on the recruitment rate over time and to describe costs of these strategies to inform future study planning and optimize resource allocation. Methods A cluster-randomized controlled trial (“AOK-Family+” study) was conducted from April 2023 to June 2024 in southwestern Germany. The intervention focused on reducing lifestyle-related risk factors (LRRFs) among pregnant women and women planning a pregnancy. Analog recruitment strategies included e.g., printed magazines, flyers and digital recruitment strategies included e.g. social media ads, influencer marketing. For 255 interested women, contact details, recruitment source, and participation status were documented and analyzed. A segmented linear regression analysis was used to identify turning points in application trends. Direct costs were calculated based on internal project budget tracking. Results The targeted sample size was not reached despite substantial investment in recruitment measures. The highest number of applications resulted from analog strategies—especially printed AOK magazines (35.5%)—followed by influencer marketing (23.6%). The segmented linear regression analysis identified three significant increases in application rates, the first coinciding with the magazine distribution and the second with the influencer marketing on Instagram. Social media marketing showed a short-lived effect, with application rates dropping immediately after posts ended. Total costs amounted to 99.334 € equaling 389,54€ per application and 534,05€ per actually enrolled participant. Conclusion Health magazines proved to be the most cost-efficient and sustainable recruitment strategy. Influencer marketing led to high reach and initial spikes in engagement but had limited long-term impact. While digital measures generated many clicks, only a small fraction translated into study participation—indicating a pronounced intention–behavior gap and procedural barriers in the enrolment process. For future studies, a mixed-methods recruitment strategy is recommended combining wide digital outreach with personalized, trust-based communication, ideally through healthcare professionals such as gynecologists and midwives, to reach women in early pregnancy and reduce participation barriers. Trial registration The German Clinical Trials Register DRKS00027804. Registered on 2022/01/12.
Stepwise regression algorithm based on the ordered pair of normalized real numbers framework
In the realm of machine learning algorithms based on the framework of ordered pair of normalized real numbers (OPNs), there is a sharp increase in the number of model feature variables when real number data are converted to the corresponding OPNs form. For real number data, the conversion results in a vast search space for the model. If these OPNs variables are entered into the machine learning model without selection, the model is prone to errors due to the large search space, and its performance will suffer dramatically. This study introduces a stepwise regression algorithm within the OPNs framework, applying binary fundamental operation units and their methods of operation from OPNs theory to stepwise regression. We select different variable screening criteria as needed, remove redundant OPNs variables, and extract significant OPNs variables to include in the regression equation. At the same time, we exclude those variables that have insignificant effects, thereby reducing unnecessary redundant information and constructing an 'optimal’ OPNs regression analysis model. We are dedicated to identifying the optimal or near-optimal OPNs variable inputs, minimizing the input of the model, and improving the performance of the algorithm. We then investigate the patterns of real number data characteristics paired as OPNs variable features, which reduces the size of input data and computational load for machine learning models based on the OPNs theory. Compared to traditional algorithms, the OPNs stepwise regression (OPNs-SR) model demonstrates more information and better algorithmic results on most datasets.
An adaptive identification method for outliers in dam deformation monitoring data based on Bayesian model selection and least trimmed squares estimation
An important technique for the quantitative analysis of dam deformation state is to establish safety monitoring models using deformation monitoring data. To address the shortcomings of conventional monitoring models, such as difficulty in selecting influencing factors and poor ability to resist the interference of outliers, this paper develops a structural safety monitoring model that can realize adaptive identification of various types of outliers in dam deformation monitoring data. The Bayesian model selection (BMS) method is first introduced to select the explanatory variables with a significant impact on the modeling process. On this basis, robust regression analysis of dam deformation monitoring data is performed by using the least trimmed squares (LTS) estimation. In particular, the recovery of clean data and the regression learning are conducted jointly. Furthermore, the double wedge plot is proposed, a graphical display which indicates outliers and potential level shifts. The engineering example demonstrates that, compared with the widely used multiple linear regression (MLR) model based on least squares (LS) fitting, the robust regression model based on BMS-LTS can not only effectively determine the key influencing factors but also adaptively identify various types of outliers in the regression. This study improves the significance of regression and increases the accuracy of prediction; thus, it has good applicability in anomaly detection of dam monitoring data and quantitative analysis of dam safety behavior.
Does social media matter for post typology? Impact of post content on Facebook and Instagram metrics
Purpose – The purpose of this paper is to measure the impact of post type (advertising, fan, events, information, and promotion) on two interaction metrics: likes and comments. The measuring involved two popular social media, Facebook and Instagram, and in business profiles of five different segments (food, hairdressing, ladies’ footwear, body design, fashion gym wear). Design/methodology/approach – The method used was multiple regression analysis with an estimator of the ordinary least squares for 1,849 posts from five different companies posted on Facebook (680 posts) and Instagram (1,169 Instagram) over an eight-month posting period. Regression analysis was used to identify the relationship between the dependent variables (likes and comments), and the independent variables (post typology, segments, week period, month, characters and hashtag). Findings – It was seen that the post types events and promotion led to a greater involvement of followers in Instagram, in particular. In Facebook, the events post type was only significant in the like’s interaction. Another finding of the research is the relevance of the food and body design segment which was significant in both virtual social media. This indicates a user preference involving their day-to-day lives, in this case, having a tattoo done or seeing a photo of a dessert. Originality/value – With the findings of this study, academics and social media managers can improve the return indicators of interactions in posts and broaden the discussion on the types of post and interaction in different virtual social media.
Novel Fuzzy Correlation Coefficient and Variable Selection Method for Fuzzy Regression Analysis Based on Distance Approach
In data analysis, analyzing the relationships between the variables such as correlation analysis and regression analysis are very important. Correlation analysis and regression analysis are not only very important in analyzing the influence relationship and causal relationship of variables but also serve as the basis for statistical analysis. Furthermore, they are essential and important as basic analysis for machine learning analysis such as deep learning. This is because in analyzing the input and output in deep learning, variables with high correlation are selected first, and in analyzing the causal relationship, it is basic to first conduct basic analysis such as regression analysis. Especially, when data are observed as fuzzy data with ambiguous information, it is difficult to propose unique methods for those analyses due to its complexity. However, the application of fuzzy theory to correlation analysis for data with such ambiguous information has not been an effective study, and several studies have been conducted in cases where the data is not general fuzzy data or interval estimation. As a result, the effectiveness of the fuzzy theory was not highlighted. In particular, the variable selection method for selecting important variables in multiple regression analysis is a very important and essential process in regression analysis. A variable that is significant in simple regression analysis may not be significant in multiple regression analysis due to its relationship with other variables. Therefore, not all variables that affect the dependent variable can be used as independent variables in multiple regression analysis. Therefore, multiple regression analysis goes through the process of excluding some variables. But until now, the process of fuzzy multiple regression analysis has not been applied without a variable selection method and the significance of important variables has not been emphasized that much. In this paper, a fuzzy correlation coefficient and multiple fuzzy regression analysis using variable section method are proposed. For this, first defuzzification and fuzzy ordering are defined. And then fuzzy correlation coefficient is proposed using L 2 distance. Next, fuzzy sum of squares are defined for F -statistics to test the significance of the regression model. Using this F -statistics, fuzzy R 2 , and fuzzy RMSE , several variable selection methods are proposed based on distance approach. For the data analysis, foreign exchange reserve data and house price of South Korea have been applied which are important indicators for economic crisis. The financial data is mostly recorded as closing values, but the closing values cannot be the representative of the given period of time. Therefore, we can deal with the financial data as fuzzy data which have some fluctuation that can be considered as vagueness that the data originally include. We have used foreign exchange reserve data and house price data with several financial variables. And the proposed fuzzy correlation coefficient and variable selection for fuzzy regression analysis are applied to these financial data.