Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
      More Filters
      Clear All
      More Filters
      Source
    • Language
2,050 result(s) for "RGB"
Sort by:
Indoor Scene Understanding with RGB-D Images: Bottom-up Segmentation, Object Detection and Semantic Segmentation
In this paper, we address the problems of contour detection, bottom-up grouping, object detection and semantic segmentation on RGB-D data. We focus on the challenging setting of cluttered indoor scenes, and evaluate our approach on the recently introduced NYU-Depth V2 (NYUD2) dataset (Silberman et al., ECCV, 2012 ). We propose algorithms for object boundary detection and hierarchical segmentation that generalize the g P b - u c m approach of Arbelaez et al. (TPAMI, 2011 ) by making effective use of depth information. We show that our system can label each contour with its type (depth, normal or albedo). We also propose a generic method for long-range amodal completion of surfaces and show its effectiveness in grouping. We train RGB-D object detectors by analyzing and computing histogram of oriented gradients on the depth image and using them with deformable part models (Felzenszwalb et al., TPAMI, 2010 ). We observe that this simple strategy for training object detectors significantly outperforms more complicated models in the literature. We then turn to the problem of semantic segmentation for which we propose an approach that classifies superpixels into the dominant object categories in the NYUD2 dataset. We design generic and class-specific features to encode the appearance and geometry of objects. We also show that additional features computed from RGB-D object detectors and scene classifiers further improves semantic segmentation accuracy. In all of these tasks, we report significant improvements over the state-of-the-art.
RGB-D face recognition using LBP with suitable feature dimension of depth image
This study proposes a robust method for the face recognition from low-resolution red, green, and blue-depth (RGB-D) cameras acquired images which have a wide range of variations in head pose, illumination, facial expression, and occlusion in some cases. The local binary pattern (LBP) of the RGB-D images with the suitable feature dimension of Depth image is employed to extract the facial features. On the basis of error correcting output codes, they are fed to multiclass support vector machines (MSVMs) for the off-line training and validation, and then the online classification. The proposed method is called as the LBP-RGB-D-MSVM with the suitable feature dimension of the depth image. The effectiveness of the proposed method is evaluated by the four databases: Indraprastha Institute of Information Technology, Delhi (IIIT-D) RGB-D, visual analysis of people (VAP) RGB-D-T, EURECOM, and the authors. In addition, an extended database merged by the first three databases is employed to compare among the proposed method and some existing two-dimensional (2D) and 3D face recognition algorithms. The proposed method possesses satisfactory performance (as high as 99.10 ± 0.52% for Rank 5 recognition rate in their database) with low computation (62 ms for feature extraction) which is desirable for real-time applications.
Video benchmarks of human action datasets: a review
Vision-based Human activity recognition is becoming a trendy area of research due to its wide application such as security and surveillance, human–computer interactions, patients monitoring system, and robotics. In the past two decades, there are several publically available human action, and activity datasets are reported based on modalities, view, actors, actions, and applications. The objective of this survey paper is to outline the different types of video datasets and highlights their merits and demerits under practical considerations. Based on the available information inside the dataset we can categorise these datasets into RGB (Red, Green, and Blue) and RGB-D(depth). The most prominent challenges involved in these datasets are occlusions, illumination variation, view variation, annotation, and fusion of modalities. The key specification of these datasets is discussed such as resolutions, frame rate, actions/actors, background, and application domain. We have also presented the state-of-the-art algorithms in a tabular form that give the best performance on such datasets. In comparison with earlier surveys, our works give a better presentation of datasets on the well-organised comparison, challenges, and latest evaluation technique on existing datasets.
CFRNet: Cross-Attention-Based Fusion and Refinement Network for Enhanced RGB-T Salient Object Detection
Existing deep learning-based RGB-T salient object detection methods often struggle with effectively fusing RGB and thermal features. Therefore, obtaining high-quality features and fully integrating these two modalities are central research focuses. We developed an illumination prior-based coefficient predictor (MICP) to determine optimal interaction weights. We then designed a saliency-guided encoder (SG Encoder) to extract multi-scale thermal features incorporating saliency information. The SG Encoder guides the extraction of thermal features by leveraging their correlation with RGB features, particularly those with strong semantic relationships to salient object detection tasks. Finally, we employed a Cross-attention-based Fusion and Refinement Module (CrossFRM) to refine the fused features. The robust thermal features help refine the spatial focus of the fused features, aligning them more closely with salient objects. Experimental results demonstrate that our proposed approach can more accurately locate salient objects, significantly improving performance compared to 11 state-of-the-art methods.
Exploring RGB+Depth Fusion for Real-Time Object Detection
In this paper, we investigate whether fusing depth information on top of normal RGB data for camera-based object detection can help to increase the performance of current state-of-the-art single-shot detection networks. Indeed, depth sensing is easily acquired using depth cameras such as a Kinect or stereo setups. We investigate the optimal manner to perform this sensor fusion with a special focus on lightweight single-pass convolutional neural network (CNN) architectures, enabling real-time processing on limited hardware. For this, we implement a network architecture allowing us to parameterize at which network layer both information sources are fused together. We performed exhaustive experiments to determine the optimal fusion point in the network, from which we can conclude that fusing towards the mid to late layers provides the best results. Our best fusion models significantly outperform the baseline RGB network in both accuracy and localization of the detections.
Enhanced ECG Signal features transformation to RGB matrix imaging for advanced deep learning classification of myocardial infarction and cardiac arrhythmia
Identifying and accurately classifying cardiac abnormalities, including myocardial infarction (MI) and cardiac arrhythmia (CA), remains a significant challenge in the field of cardiology, largely due to the limitations inherent in traditional ECG signal analysis techniques. This paper presents an innovative method aimed at addressing this challenge. By implementing a novel transformation technique, we map temporal, frequency-based, statistical, and spatial features of ECG signals onto the R, G, and B channels of an RGB image. This conversion process results in a feature-rich representation of the ECG signal, significantly enhancing its clinical relevance and thus maximizing classification accuracy. Utilizing an adaptive RGB-ResNet inception architecture, our approach achieves remarkable average accuracies of 99.25% for myocardial infarction and 99.21% for cardiac arrhythmia. These figures underscore the robustness of our method and highlight its significant potential to advance cardiology diagnostics through the application of advanced image analysis techniques.
Computation of Gait Parameters in Post Stroke and Parkinson’s Disease: A Comparative Study Using RGB-D Sensors and Optoelectronic Systems
The accurate and reliable assessment of gait parameters is assuming an important role, especially in the perspective of designing new therapeutic and rehabilitation strategies for the remote follow-up of people affected by disabling neurological diseases, including Parkinson’s disease and post-stroke injuries, in particular considering how gait represents a fundamental motor activity for the autonomy, domestic or otherwise, and the health of neurological patients. To this end, the study presents an easy-to-use and non-invasive solution, based on a single RGB-D sensor, to estimate specific features of gait patterns on a reduced walking path compatible with the available spaces in domestic settings. Traditional spatio-temporal parameters and features linked to dynamic instability during walking are estimated on a cohort of ten parkinsonian and eleven post-stroke subjects using a custom-written software that works on the result of a body-tracking algorithm. Then, they are compared with the “gold standard” 3D instrumented gait analysis system. The statistical analysis confirms no statistical difference between the two systems. Data also indicate that the RGB-D system is able to estimate features of gait patterns in pathological individuals and differences between them in line with other studies. Although they are preliminary, the results suggest that this solution could be clinically helpful in evolutionary disease monitoring, especially in domestic and unsupervised environments where traditional gait analysis is not usable.
Autonomous Exploration of Unknown Indoor Environments for High-Quality Mapping Using Feature-Based RGB-D SLAM
Simultaneous localization and mapping (SLAM) system-based indoor mapping using autonomous mobile robots in unknown environments is crucial for many applications, such as rescue scenarios, utility tunnel monitoring, and indoor 3D modeling. Researchers have proposed various strategies to obtain full coverage while minimizing exploration time; however, mapping quality factors have not been considered. In fact, mapping quality plays a pivotal role in 3D modeling, especially when using low-cost sensors in challenging indoor scenarios. This study proposes a novel exploration algorithm to simultaneously optimize exploration time and mapping quality using a low-cost RGB-D camera. Feature-based RGB-D SLAM is utilized due to its various advantages, such as low computational cost and dense real-time reconstruction ability. Subsequently, our novel exploration strategies consider the mapping quality factors of the RGB-D SLAM system. Exploration time optimization factors are also considered to set a new optimum goal. Furthermore, a Voronoi path planner is adopted for reliable, maximal obstacle clearance and fixed paths. According to the texture level, three exploration strategies are evaluated in three real-world environments. We achieve a significant enhancement in mapping quality and exploration time using our proposed exploration strategies compared to the baseline frontier-based exploration, particularly in a low-texture environment.
Minimum Spanning Tree Image Segmentation Model Based on New Weights
The quality of image segmentation results directly affects subsequent image processing and its application, therefore image segmentation is the most crucial step in image processing and recognition. A minimum spanning tree image segmentation method using new weights is proposed to address the issues of low efficiency and low segmentation accuracy in traditional image segmentation methods. The study first proposed a minimum spanning tree construction method based on color vector angle distance color difference measurement, which compensates for the non-uniformity of color space through dynamic weight adjustment, and improves on the fast multi spanning tree segmentation algorithm with adaptive threshold. The proposed decomposition method preprocesses the image to reduce the image nodes. The experiment showed that the research method effectively improved color difference measurement accuracy. The segmentation accuracy, over segmentation rate, and under segmentation rate obtained by the research method outperformed other segmentation methods in terms of average values, with an average of 0.984, 0.059, and 0.023, respectively. Compared to other minimum spanning tree segmentation algorithms, the research method improved the average segmentation time by 0.46 seconds, 0.49 seconds, 2.04 seconds, and 3.79 seconds, respectively. The segmentation algorithm studied in this study has good segmentation performance and improves the efficiency of image segmentation, which has certain practical application value in various image segmentation fields.
Learnable Depth-Sensitive Attention for Deep RGB-D Saliency Detection with Multi-modal Fusion Architecture Search
RGB-D salient object detection (SOD) is usually formulated as a problem of classification or regression over two modalities, i.e. , RGB and depth. Hence, effective RGB-D feature modeling and multi-modal feature fusion both play a vital role in RGB-D SOD. In this paper, we propose a depth-sensitive RGB feature modeling scheme using the depth-wise geometric prior of salient objects. In principle, the feature modeling scheme is carried out in a Depth-Sensitive Attention Module (DSAM), which leads to the RGB feature enhancement as well as the background distraction reduction by capturing the depth geometry prior. Furthermore, we extend and enhance the original DSAM to DSAMv2 by proposing a novel Depth Attention Generation Module (DAGM) to generate learnable depth attention maps for more robust depth-sensitive RGB feature extraction. Moreover, to perform effective multi-modal feature fusion, we further present an automatic neural architecture search approach for RGB-D SOD, which does well in finding out a feasible architecture from our specially designed multi-modal multi-scale search space. Extensive experiments on nine standard benchmarks have demonstrated the effectiveness of the proposed approach against the state-of-the-art. We name the enhanced learnable Depth-Sensitive Attention and Automatic multi-modal Fusion framework DSA2Fv2.