Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
      More Filters
      Clear All
      More Filters
      Source
    • Language
93,625 result(s) for "classification algorithms"
Sort by:
Supervised machine learning algorithms for predicting student dropout and academic success: a comparative study
Utilizing a dataset sourced from a higher education institution, this study aims to assess the efficacy of diverse machine learning algorithms in predicting student dropout and academic success. Our focus was on algorithms capable of effectively handling imbalanced data. To tackle class imbalance, we employed the SMOTE resampling technique. We applied a range of algorithms, including Decision Tree (DT), Support Vector Machine (SVM), Random Forest (RF), as well as boosting algorithms such as Gradient Boosting (GB), Extreme Gradient Boosting (XGBoost), CatBoost (CB), and Light Gradient Boosting Machine (LB). To enhance the models' performance, we conducted hyperparameter tuning using Optuna. Additionally, we employed the Isolation Forest (IF) method to identify outliers or anomalies within the dataset. Notably, our findings indicate that boosting algorithms, particularly LightGBM and CatBoost with Optuna, outperformed traditional classification methods. Our study's generalizability to other contexts is constrained due to its reliance on a single dataset, with inherent limitations. Nevertheless, this research provides valuable insights into the effectiveness of various machine learning algorithms for predicting student dropout and academic success. By benchmarking these algorithms, our project offers guidance to both researchers and practitioners in their choice of suitable approaches for similar predictive tasks.
Analyzing Student Performance in Programming Education Using Classification Techniques
In this research, we aggregated students log data such as Class Test Score (CTS), Assignment Completed (ASC), Class Lab Work (CLW) and Class Attendance (CATT) from the Department of Mathematics, Computer Science Unit, Usmanu Danfodiyo University, Sokoto, Nigeria. Similarly, we employed data mining techniques such as ID3 & J48 Decision Tree Algorithms to analyze these data. We compared these algorithms on 239 classification instances. The experimental results show that the J48 algorithm has higher accuracy in the classification task compared to the ID3 algorithm. The important feature attributes such as Information Gain and Gain Ratio feature evaluators were also compared. Both the methods applied were able to rank search method and the experimental results confirmed that the two methods derived the same set of attributes with a slight deviation in the ranking. From the results analyzed, we discovered that 67.36 percent failed the course titled Introduction to Computer Programming, while 32.64 percent passed the course. Since the CATT has the highest gain value from our analysis; we concluded that it is largely responsible for the success or failure of the students.
Assessment of Machine Learning Techniques for Oil Rig Classification in C-Band SAR Images
This article aims at performing maritime target classification in SAR images using machine learning (ML) and deep learning (DL) techniques. In particular, the targets of interest are oil platforms and ships located in the Campos Basin, Brazil. Two convolutional neural networks (CNNs), VGG-16 and VGG-19, were used for attribute extraction. The logistic regression (LR), random forest (RF), support vector machine (SVM), k-nearest neighbours (kNN), decision tree (DT), naive Bayes (NB), neural networks (NET), and AdaBoost (ADBST) schemes were considered for classification. The target classification methods were evaluated using polarimetric images obtained from the C-band synthetic aperture radar (SAR) system Sentinel-1. Classifiers are assessed by the accuracy indicator. The LR, SVM, NET, and stacking results indicate better performance, with accuracy ranging from 84.1% to 85.5%. The Kruskal–Wallis test shows a significant difference with the tested classifier, indicating that some classifiers present different accuracy results. The optimizations provide results with more significant accuracy gains, making them competitive with those shown in the literature. There is no exact combination of methods for SAR image classification that will always guarantee the best accuracy. The optimizations performed in this article were for the specific data set of the Campos Basin, and results may change depending on the data set format and the number of images.
Hybrid feature selection and classification model using high-dimensional data based on a metaheuristic algorithm for brain cancer diagnosis
Cancer is caused by somatic mutations, a dreadful disease that impacts individuals everywhere. Classifying gene expression data is essential for disease diagnosis and distinguishing tumor types. However, small sample sizes, numerous features, and noise make this task particularly challenging. This is especially true when performing feature selection on high-dimensional microarray data. It is critical to select the most pertinent and valuable genes from microarray data to identify prospective biomarkers or gain insight into the fundamental mechanisms of cancer. This study introduces a novel hybrid model that combines feature selection and classification to identify the most significant and informative features from microarray data associated with brain cancer. The research employs the GSE50161 dataset obtained from the Curated Microarray Database (CuMiDa), comprising 130 samples classified into five distinct categories with 54,676 genomes examined. We first applied mRMR to reduce dimensionality by removing redundant features, followed by HHO to refine the feature subset for optimal classification performance. To improve the performance of our model in classifying brain cancer microarray data, we utilized three metaheuristic algorithms: Differential Evolution (DE), Harris Hawks Optimization (HHO), and Particle Swarm Optimization (PSO). The hyperparameters “C” and “sigma” of the support vector machine (SVM) were optimized using these algorithms. The experimental results indicate that the suggested framework improves the capacity to differentiate between benign and malignant tissues with reduced time and dimensionality requirements. Furthermore, the genes selected for the dataset on brain cancer have undergone biological interpretation. This process is consistent with the findings of relevant scientific inquiries and significantly influences patients’ prognoses.
Vehicle maintenance management based on machine learning in agricultural tractor engines
The objective of this work is to use the autonomous learning methodology as a tool in vehicle maintenance management. In obtaining data, faults in the fuel supply system have been simulated, causing anomalies in the combustion process that are easily detectable by vibrations obtained from a sensor in the engine of an agricultural tractor. To train the classification algorithm, 4 engine states were used: BE (optimal state), MEF1, MEF2, MEF3 (simulated failures). The applied autonomous learning is of the supervised type, where the samples were initially characterized and labeled to create a database for the execution of the training. The results show that the training carried out within the classification algorithm has an efficiency greater than 90%, which indicates that the method used is applicable in the management of vehicle maintenance to predict failures in engine operation.
Comparison of Random Forest, k-Nearest Neighbor, and Support Vector Machine Classifiers for Land Cover Classification Using Sentinel-2 Imagery
In previous classification studies, three non-parametric classifiers, Random Forest (RF), k-Nearest Neighbor (kNN), and Support Vector Machine (SVM), were reported as the foremost classifiers at producing high accuracies. However, only a few studies have compared the performances of these classifiers with different training sample sizes for the same remote sensing images, particularly the Sentinel-2 Multispectral Imager (MSI). In this study, we examined and compared the performances of the RF, kNN, and SVM classifiers for land use/cover classification using Sentinel-2 image data. An area of 30 × 30 km2 within the Red River Delta of Vietnam with six land use/cover types was classified using 14 different training sample sizes, including balanced and imbalanced, from 50 to over 1250 pixels/class. All classification results showed a high overall accuracy (OA) ranging from 90% to 95%. Among the three classifiers and 14 sub-datasets, SVM produced the highest OA with the least sensitivity to the training sample sizes, followed consecutively by RF and kNN. In relation to the sample size, all three classifiers showed a similar and high OA (over 93.85%) when the training sample size was large enough, i.e., greater than 750 pixels/class or representing an area of approximately 0.25% of the total study area. The high accuracy was achieved with both imbalanced and balanced datasets.
Unsupervised classification of scattering behaviour using hybrid-polarimetry
This study presents an unsupervised algorithm for classification of scattering behaviour using hybrid-polarimetric (hybrid-Pol) data. The authors present a maximum likelihood estimation-based unsupervised land cover classification algorithms for hybrid-PolSAR image. This classification technique follows from the m − δ decomposition of hybrid-Pol images. Introduction of a statistical treatment is the major contribution of the current algorithm. Performance of the hybrid-Pol algorithms have been assessed with respect to Freeman–Durden decomposition of fully polarimetric SAR data. The authors have demonstrated, using two different datasets, that proposed algorithm not only gives better overall classification performance, it is also able to classify all the three major types of scattering mechanisms, whereas the existing hybrid-PolSAR classification algorithms mostly fail to classify one of the scattering types.
Recent advances in decision trees: an updated survey
Decision Trees (DTs) are predictive models in supervised learning, known not only for their unquestionable utility in a wide range of applications but also for their interpretability and robustness. Research on the subject is still going strong after almost 60 years since its original inception, and in the last decade, several researchers have tackled key matters in the field. Although many great surveys have been published in the past, there is a gap since none covers the last decade of the field as a whole. This paper proposes a review of the main recent advances in DT research, focusing on three major goals of a predictive learner: issues regarding the fitting of training data, generalization, and interpretability. Moreover, by organizing several topics that have been previously analyzed in isolation, this survey attempts to provide an overview of the field, its key concerns, and future trends, serving as a good entry point for both researchers and newcomers to the machine learning community.
Analysis and Text Classification of Privacy Policies From Rogue and Top-100 Fortune Global Companies
In the present article, the authors investigate to what extent supervised binary classification can be used to distinguish between legitimate and rogue privacy policies posted on web pages. 15 classification algorithms are evaluated using a data set that consists of 100 privacy policies from legitimate websites (belonging to companies that top the Fortune Global 500 list) as well as 67 policies from rogue websites. A manual analysis of all policy content was performed and clear statistical differences in terms of both length and adherence to seven general privacy principles are found. Privacy policies from legitimate companies have a 98% adherence to the seven privacy principles, which is significantly higher than the 45% associated with rogue companies. Out of the 15 evaluated classification algorithms, Naïve Bayes Multinomial is the most suitable candidate to solve the problem at hand. Its models show the best performance, with an AUC measure of 0.90 (0.08), which outperforms most of the other candidates in the statistical tests used.
Decision trees: a recent overview
Decision tree techniques have been widely used to build classification models as such models closely resemble human reasoning and are easy to understand. This paper describes basic decision tree issues and current research points. Of course, a single article cannot be a complete review of all algorithms (also known induction classification trees), yet we hope that the references cited will cover the major theoretical issues, guiding the researcher in interesting research directions and suggesting possible bias combinations that have yet to be explored.