Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
2,050
result(s) for
"RGB"
Sort by:
Indoor Scene Understanding with RGB-D Images: Bottom-up Segmentation, Object Detection and Semantic Segmentation
by
Malik, Jitendra
,
Arbeláez, Pablo
,
Girshick, Ross
in
Algorithms
,
Analysis
,
Artificial Intelligence
2015
In this paper, we address the problems of contour detection, bottom-up grouping, object detection and semantic segmentation on RGB-D data. We focus on the challenging setting of cluttered indoor scenes, and evaluate our approach on the recently introduced NYU-Depth V2 (NYUD2) dataset (Silberman et al., ECCV,
2012
). We propose algorithms for object boundary detection and hierarchical segmentation that generalize the
g
P
b
-
u
c
m
approach of Arbelaez et al. (TPAMI,
2011
) by making effective use of depth information. We show that our system can label each contour with its type (depth, normal or albedo). We also propose a generic method for long-range amodal completion of surfaces and show its effectiveness in grouping. We train RGB-D object detectors by analyzing and computing histogram of oriented gradients on the depth image and using them with deformable part models (Felzenszwalb et al., TPAMI,
2010
). We observe that this simple strategy for training object detectors significantly outperforms more complicated models in the literature. We then turn to the problem of semantic segmentation for which we propose an approach that classifies superpixels into the dominant object categories in the NYUD2 dataset. We design generic and class-specific features to encode the appearance and geometry of objects. We also show that additional features computed from RGB-D object detectors and scene classifiers further improves semantic segmentation accuracy. In all of these tasks, we report significant improvements over the state-of-the-art.
Journal Article
RGB-D face recognition using LBP with suitable feature dimension of depth image
2019
This study proposes a robust method for the face recognition from low-resolution red, green, and blue-depth (RGB-D) cameras acquired images which have a wide range of variations in head pose, illumination, facial expression, and occlusion in some cases. The local binary pattern (LBP) of the RGB-D images with the suitable feature dimension of Depth image is employed to extract the facial features. On the basis of error correcting output codes, they are fed to multiclass support vector machines (MSVMs) for the off-line training and validation, and then the online classification. The proposed method is called as the LBP-RGB-D-MSVM with the suitable feature dimension of the depth image. The effectiveness of the proposed method is evaluated by the four databases: Indraprastha Institute of Information Technology, Delhi (IIIT-D) RGB-D, visual analysis of people (VAP) RGB-D-T, EURECOM, and the authors. In addition, an extended database merged by the first three databases is employed to compare among the proposed method and some existing two-dimensional (2D) and 3D face recognition algorithms. The proposed method possesses satisfactory performance (as high as 99.10 ± 0.52% for Rank 5 recognition rate in their database) with low computation (62 ms for feature extraction) which is desirable for real-time applications.
Journal Article
Video benchmarks of human action datasets: a review
2019
Vision-based Human activity recognition is becoming a trendy area of research due to its wide application such as security and surveillance, human–computer interactions, patients monitoring system, and robotics. In the past two decades, there are several publically available human action, and activity datasets are reported based on modalities, view, actors, actions, and applications. The objective of this survey paper is to outline the different types of video datasets and highlights their merits and demerits under practical considerations. Based on the available information inside the dataset we can categorise these datasets into RGB (Red, Green, and Blue) and RGB-D(depth). The most prominent challenges involved in these datasets are occlusions, illumination variation, view variation, annotation, and fusion of modalities. The key specification of these datasets is discussed such as resolutions, frame rate, actions/actors, background, and application domain. We have also presented the state-of-the-art algorithms in a tabular form that give the best performance on such datasets. In comparison with earlier surveys, our works give a better presentation of datasets on the well-organised comparison, challenges, and latest evaluation technique on existing datasets.
Journal Article
CFRNet: Cross-Attention-Based Fusion and Refinement Network for Enhanced RGB-T Salient Object Detection
2024
Existing deep learning-based RGB-T salient object detection methods often struggle with effectively fusing RGB and thermal features. Therefore, obtaining high-quality features and fully integrating these two modalities are central research focuses. We developed an illumination prior-based coefficient predictor (MICP) to determine optimal interaction weights. We then designed a saliency-guided encoder (SG Encoder) to extract multi-scale thermal features incorporating saliency information. The SG Encoder guides the extraction of thermal features by leveraging their correlation with RGB features, particularly those with strong semantic relationships to salient object detection tasks. Finally, we employed a Cross-attention-based Fusion and Refinement Module (CrossFRM) to refine the fused features. The robust thermal features help refine the spatial focus of the fused features, aligning them more closely with salient objects. Experimental results demonstrate that our proposed approach can more accurately locate salient objects, significantly improving performance compared to 11 state-of-the-art methods.
Journal Article
Exploring RGB+Depth Fusion for Real-Time Object Detection
by
Van Beeck, Kristof
,
Goedemé, Toon
,
Ophoff, Tanguy
in
Depth
,
Neural Networks
,
Object detection
2019
In this paper, we investigate whether fusing depth information on top of normal RGB data for camera-based object detection can help to increase the performance of current state-of-the-art single-shot detection networks. Indeed, depth sensing is easily acquired using depth cameras such as a Kinect or stereo setups. We investigate the optimal manner to perform this sensor fusion with a special focus on lightweight single-pass convolutional neural network (CNN) architectures, enabling real-time processing on limited hardware. For this, we implement a network architecture allowing us to parameterize at which network layer both information sources are fused together. We performed exhaustive experiments to determine the optimal fusion point in the network, from which we can conclude that fusing towards the mid to late layers provides the best results. Our best fusion models significantly outperform the baseline RGB network in both accuracy and localization of the detections.
Journal Article
Enhanced ECG Signal features transformation to RGB matrix imaging for advanced deep learning classification of myocardial infarction and cardiac arrhythmia
2025
Identifying and accurately classifying cardiac abnormalities, including myocardial infarction (MI) and cardiac arrhythmia (CA), remains a significant challenge in the field of cardiology, largely due to the limitations inherent in traditional ECG signal analysis techniques. This paper presents an innovative method aimed at addressing this challenge. By implementing a novel transformation technique, we map temporal, frequency-based, statistical, and spatial features of ECG signals onto the R, G, and B channels of an RGB image. This conversion process results in a feature-rich representation of the ECG signal, significantly enhancing its clinical relevance and thus maximizing classification accuracy. Utilizing an adaptive RGB-ResNet inception architecture, our approach achieves remarkable average accuracies of 99.25% for myocardial infarction and 99.21% for cardiac arrhythmia. These figures underscore the robustness of our method and highlight its significant potential to advance cardiology diagnostics through the application of advanced image analysis techniques.
Journal Article
Computation of Gait Parameters in Post Stroke and Parkinson’s Disease: A Comparative Study Using RGB-D Sensors and Optoelectronic Systems
2022
The accurate and reliable assessment of gait parameters is assuming an important role, especially in the perspective of designing new therapeutic and rehabilitation strategies for the remote follow-up of people affected by disabling neurological diseases, including Parkinson’s disease and post-stroke injuries, in particular considering how gait represents a fundamental motor activity for the autonomy, domestic or otherwise, and the health of neurological patients. To this end, the study presents an easy-to-use and non-invasive solution, based on a single RGB-D sensor, to estimate specific features of gait patterns on a reduced walking path compatible with the available spaces in domestic settings. Traditional spatio-temporal parameters and features linked to dynamic instability during walking are estimated on a cohort of ten parkinsonian and eleven post-stroke subjects using a custom-written software that works on the result of a body-tracking algorithm. Then, they are compared with the “gold standard” 3D instrumented gait analysis system. The statistical analysis confirms no statistical difference between the two systems. Data also indicate that the RGB-D system is able to estimate features of gait patterns in pathological individuals and differences between them in line with other studies. Although they are preliminary, the results suggest that this solution could be clinically helpful in evolutionary disease monitoring, especially in domestic and unsupervised environments where traditional gait analysis is not usable.
Journal Article
Autonomous Exploration of Unknown Indoor Environments for High-Quality Mapping Using Feature-Based RGB-D SLAM
2022
Simultaneous localization and mapping (SLAM) system-based indoor mapping using autonomous mobile robots in unknown environments is crucial for many applications, such as rescue scenarios, utility tunnel monitoring, and indoor 3D modeling. Researchers have proposed various strategies to obtain full coverage while minimizing exploration time; however, mapping quality factors have not been considered. In fact, mapping quality plays a pivotal role in 3D modeling, especially when using low-cost sensors in challenging indoor scenarios. This study proposes a novel exploration algorithm to simultaneously optimize exploration time and mapping quality using a low-cost RGB-D camera. Feature-based RGB-D SLAM is utilized due to its various advantages, such as low computational cost and dense real-time reconstruction ability. Subsequently, our novel exploration strategies consider the mapping quality factors of the RGB-D SLAM system. Exploration time optimization factors are also considered to set a new optimum goal. Furthermore, a Voronoi path planner is adopted for reliable, maximal obstacle clearance and fixed paths. According to the texture level, three exploration strategies are evaluated in three real-world environments. We achieve a significant enhancement in mapping quality and exploration time using our proposed exploration strategies compared to the baseline frontier-based exploration, particularly in a low-texture environment.
Journal Article
Minimum Spanning Tree Image Segmentation Model Based on New Weights
by
Li, Hong
2024
The quality of image segmentation results directly affects subsequent image processing and its application, therefore image segmentation is the most crucial step in image processing and recognition. A minimum spanning tree image segmentation method using new weights is proposed to address the issues of low efficiency and low segmentation accuracy in traditional image segmentation methods. The study first proposed a minimum spanning tree construction method based on color vector angle distance color difference measurement, which compensates for the non-uniformity of color space through dynamic weight adjustment, and improves on the fast multi spanning tree segmentation algorithm with adaptive threshold. The proposed decomposition method preprocesses the image to reduce the image nodes. The experiment showed that the research method effectively improved color difference measurement accuracy. The segmentation accuracy, over segmentation rate, and under segmentation rate obtained by the research method outperformed other segmentation methods in terms of average values, with an average of 0.984, 0.059, and 0.023, respectively. Compared to other minimum spanning tree segmentation algorithms, the research method improved the average segmentation time by 0.46 seconds, 0.49 seconds, 2.04 seconds, and 3.79 seconds, respectively. The segmentation algorithm studied in this study has good segmentation performance and improves the efficiency of image segmentation, which has certain practical application value in various image segmentation fields.
Journal Article
Learnable Depth-Sensitive Attention for Deep RGB-D Saliency Detection with Multi-modal Fusion Architecture Search
2022
RGB-D salient object detection (SOD) is usually formulated as a problem of classification or regression over two modalities, i.e. , RGB and depth. Hence, effective RGB-D feature modeling and multi-modal feature fusion both play a vital role in RGB-D SOD. In this paper, we propose a depth-sensitive RGB feature modeling scheme using the depth-wise geometric prior of salient objects. In principle, the feature modeling scheme is carried out in a Depth-Sensitive Attention Module (DSAM), which leads to the RGB feature enhancement as well as the background distraction reduction by capturing the depth geometry prior. Furthermore, we extend and enhance the original DSAM to DSAMv2 by proposing a novel Depth Attention Generation Module (DAGM) to generate learnable depth attention maps for more robust depth-sensitive RGB feature extraction. Moreover, to perform effective multi-modal feature fusion, we further present an automatic neural architecture search approach for RGB-D SOD, which does well in finding out a feasible architecture from our specially designed multi-modal multi-scale search space. Extensive experiments on nine standard benchmarks have demonstrated the effectiveness of the proposed approach against the state-of-the-art. We name the enhanced learnable Depth-Sensitive Attention and Automatic multi-modal Fusion framework DSA2Fv2.
Journal Article