Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
506
result(s) for
"depth fusion"
Sort by:
Pose-Perceptive Convolution: Learning Geometry-Adaptive Receptive Fields for Robust 6D Pose Estimation
2026
6D object pose estimation is crucial for applications such as robotic manipulation and augmented reality, yet it remains highly challenging when dealing with objects of significantly different aspect ratios or the drastic appearance variations of a single object caused by pose changes. Most existing methods focus on designing more complex backend fusion modules, while largely overlooking a fundamental problem at the feature extraction frontend: the geometric mismatch between the fixed, square receptive fields of standard convolutions and the varied projected morphologies of objects. This mismatch, along with noise in fused features and ambiguity in regression, limits the performance ceiling of current methods. To this end, this paper proposes a novel Pose-Perceptive Convolution (PPC) and constructs a new Pose-Perceptive Fusion Network (PPF-Net). Its core component, the Pose-Perceptive Convolution, fundamentally resolves the aforementioned geometric mismatch by dynamically adapting the shape and sampling density of its receptive field. Experiments on four benchmarks show that PPF-Net improves the VSD score by 19.4% over FFB6D on MP6D, and achieves 96.7% ADD-S on YCB-Video, approaching state-of-the-art accuracy. Crucially, these gains are realized with minimal computational overhead, avoiding the heavy latency of backend-intensive approaches. This validates that frontend feature extraction is an efficient strategy for robust 6D pose estimation.
Journal Article
RD-GuideNet: A Depth-Guided Framework for Robust Detection, Segmentation, and Temporal Tracking of White Button Mushrooms
2026
Mushroom farms in the United States continue to face persistent labor shortages, especially during the harvesting of white button mushrooms (Agaricus bisporus) which requires selective picking by skilled workers. This study addresses this challenge by developing a depth-guided computer vision framework for automated mushroom detection, segmentation, and tracking to support timely harvest decisions, providing the foundation needed to support selective and timely robotic harvesting. The specific objectives of the study were to (1) develop a novel image-processing algorithm (RD-GuideNet) that integrates RGB and depth images for accurate detection and segmentation of mushrooms; (2) implement a custom depth-guided tracking algorithm to preserve mushroom identities across sequential frames; (3) compare the performance of RD-GuideNet against state-of-the-art deep learning models, YOLOv8 and YOLOv11, focusing on segmentation and tracking accuracies. The proposed RD-GuideNet achieved an F1-score of 0.93 for segmentation, outperforming YOLOv8 (0.88) and YOLOv11 (0.86), and produced sharper, more geometrically consistent boundaries that closely followed true mushroom cap contours. Its tracking consistency reached 92.7%, compared to YOLOv8 (95.3%) and YOLOv11 (94.6%). Although slightly lower, RD-GuideNet maintained high temporal consistency across dense mushroom beds. These results suggest that depth-based geometric reasoning and deep learning approaches exhibit complementary strengths in dense production scenes. Combining the two may further enhance detection reliability and shape fidelity, supporting high-precision perception for autonomous mushroom harvesting. A comprehensive quantitative evaluation of such a hybrid framework will be investigated in future work.
Journal Article
Efficient Depth Fusion Transformer for Aerial Image Semantic Segmentation
2022
Taking depth into consideration has been proven to improve the performance of semantic segmentation through providing additional geometry information. Most existing works adopt a two-stream network, extracting features from color images and depth images separately using two branches of the same structure, which suffer from high memory and computation costs. We find that depth features acquired by simple downsampling can also play a complementary part in the semantic segmentation task, sometimes even better than the two-stream scheme with the same two branches. In this paper, a novel and efficient depth fusion transformer network for aerial image segmentation is proposed. The presented network utilizes patch merging to downsample depth input and a depth-aware self-attention (DSA) module is designed to mitigate the gap caused by difference between two branches and two modalities. Concretely, the DSA fuses depth features and color features by computing depth similarity and impact on self-attention map calculated by color feature. Extensive experiments on the ISPRS 2D semantic segmentation dataset validate the efficiency and effectiveness of our method. With nearly half the parameters of traditional two-stream scheme, our method acquires 83.82% mIoU on Vaihingen dataset outperforming other state-of-the-art methods and 87.43% mIoU on Potsdam dataset comparable to the state-of-the-art.
Journal Article
DFusion: Denoised TSDF Fusion of Multiple Depth Maps with Sensor Pose Noises
2022
The truncated signed distance function (TSDF) fusion is one of the key operations in the 3D reconstruction process. However, existing TSDF fusion methods usually suffer from the inevitable sensor noises. In this paper, we propose a new TSDF fusion network, named DFusion, to minimize the influences from the two most common sensor noises, i.e., depth noises and pose noises. To the best of our knowledge, this is the first depth fusion for resolving both depth noises and pose noises. DFusion consists of a fusion module, which fuses depth maps together and generates a TSDF volume, as well as the following denoising module, which takes the TSDF volume as the input and removes both depth noises and pose noises. To utilize the 3D structural information of the TSDF volume, 3D convolutional layers are used in the encoder and decoder parts of the denoising module. In addition, a specially-designed loss function is adopted to improve the fusion performance in object and surface regions. The experiments are conducted on a synthetic dataset as well as a real-scene dataset. The results prove that our method outperforms existing methods.
Journal Article
On Segment-Aware Monocular Depth Estimation Using Vision Transformers
by
Mitianoudis, Nikolaos
,
Arampatzakis, Vasileios
,
Pavlidis, George
in
Accuracy
,
Architecture
,
Boundaries
2026
Monocular Depth Estimation (MDE) infers per-pixel scene geometry from a single RGB image. Despite recent progress, global MDE models often blur depth discontinuities at object boundaries and fail to capture object-level structure. Segment-aware depth estimation addresses this limitation by exploiting semantic segmentation to decompose depth prediction into simpler, class-specific subproblems. In this work, we study semantic-aware MDE in a multi-branch design where each semantic class is handled by a lightweight Vision Transformer (ViT) branch that predicts dense depth for its class while suppressing interference from other regions. We further examine fusion strategies that merge the branch outputs into a single prediction: (i) a learnable cross-attention fusion module that predicts depth from the stack of per-class proposals and masks, and (ii) a parameter-free stitched summation that sums mask-gated outputs. The proposed architecture is simple, scalable, end-to-end trainable, and compatible with arbitrary transformer backbones. Experiments on Virtual KITTI 2, where ground-truth depth and semantic labels are available, show that segment-aware modeling produces sharper depth boundaries and improves standard error metrics compared to a single-branch baseline (AbsRel 0.243→0.152; RMSE 11.952→9.101). Finally, we find that the parameter-free summation matches, and in most cases improves upon, the accuracy of learned fusion while adding no computational overhead.
Journal Article
Obstacle detection based on depth fusion of lidar and radar in challenging conditions
2021
PurposeTo the industrial application of intelligent and connected vehicles (ICVs), the robustness and accuracy of environmental perception are critical in challenging conditions. However, the accuracy of perception is closely related to the performance of sensors configured on the vehicle. To enhance sensors’ performance further to improve the accuracy of environmental perception, this paper aims to introduce an obstacle detection method based on the depth fusion of lidar and radar in challenging conditions, which could reduce the false rate resulting from sensors’ misdetection.Design/methodology/approachFirstly, a multi-layer self-calibration method is proposed based on the spatial and temporal relationships. Next, a depth fusion model is proposed to improve the performance of obstacle detection in challenging conditions. Finally, the study tests are carried out in challenging conditions, including straight unstructured road, unstructured road with rough surface and unstructured road with heavy dust or mist.FindingsThe experimental tests in challenging conditions demonstrate that the depth fusion model, comparing with the use of a single sensor, can filter out the false alarm of radar and point clouds of dust or mist received by lidar. So, the accuracy of objects detection is also improved under challenging conditions.Originality/valueA multi-layer self-calibration method is conducive to improve the accuracy of the calibration and reduce the workload of manual calibration. Next, a depth fusion model based on lidar and radar can effectively get high precision by way of filtering out the false alarm of radar and point clouds of dust or mist received by lidar, which could improve ICVs’ performance in challenging conditions.
Journal Article
Towards Neural Multi View 3D Reconstruction from RGB-D Data
2025
AbstractMulti-view stereo (MVS) networks have recently achieved remarkable progress in dense 3D reconstruction, yet they remain fundamentally limited by reliance on photometric cues. As a result, current methods fail in textureless, reflective, or non-Lambertian regions. At the same time, commodity time-of-flight (ToF) sensors provide geometric depth information that is complementary but low-resolution and noisy. In this work study a possibility to use 3D features extracted from depth data to overcome MVS limitations. For this we develop RGB-D MVSNet, an end-to-end architecture that integrates a depth-fusion encoder with a modern learning-based MVS backbone. Our method constructs a unified feature volume from both photometric and geometric features, which is then fused and regularized in a with common decoder. We evaluate the approach on the challenging Sk3D dataset containing synchronized RGB, ToF depth, and high-quality structured-light scans. Experiments demonstrate that our method improves accuracy and completeness metrics over the RGB-only baseline and achieves some qualitative improvements in reconstructing textureless and glossy regions. Additional experiments with high-quality depth input show that the method is capable of eliminating typical artifacts with better input depth quality. These results indicate that integrating geometric cues into MVS pipelines is a promising direction towards more robust, generalizable 3D reconstruction.
Journal Article
GNSS-R-Based Snow Water Equivalent Estimation with Empirical Modeling and Enhanced SNR-Based Snow Depth Estimation
2020
Snow depth and snow water equivalent (SWE) are two parameters for measuring snowfall. By exploiting the Global Navigation Satellite System reflectometry (GNSS-R) technique and thousands of existing GNSS Continuous Operating Reference Stations (CORS) deployed in the cryosphere, it is possible to improve the temporal and spatial resolutions of the SWE measurement. In this paper, a fusion model for combining multi-satellite SNR (Signal to Noise Ratio) snow depth estimations is proposed, which uses peak spectral powers associated with each of the snow depth estimations. To simplify the estimation of SWE, the complete snowfall period over a winter season is split into snow accumulation, transition, and melting period in accordance with the variation characteristics of snow depth and SWE. By extensively using in situ snow depth and SWE observations recorded by snow telemetry network (SNOTEL) and regression analysis, three empirical models are developed to describe the relationship between snow depth and SWE for the three periods, respectively. Based on the snow depth fusion model and the SWE empirical models, an SWE estimation algorithm is proposed. Three data sets recorded in different environments are used to test the proposed method. The results demonstrate that there exists good agreement between the in situ SWE measurements and the SWE estimates produced by the proposed method; the root-mean-square error of SWE estimations is smaller than 6 cm when the SWE is up to 80 cm.
Journal Article
Uncertainty-Guided Depth Fusion from Multi-View Satellite Images to Improve the Accuracy in Large-Scale DSM Generation
2022
The generation of digital surface models (DSMs) from multi-view high-resolution (VHR) satellite imagery has recently received a great attention due to the increasing availability of such space-based datasets. Existing production-level pipelines primarily adopt a multi-view stereo (MVS) paradigm, which exploit the statistical depth fusion of multiple DSMs generated from individual stereo pairs. To make this process scalable, these depth fusion methods often adopt simple approaches such as the median filter or its variants, which are efficient in computation but lack the flexibility to adapt to heterogenous information of individual pixels. These simple fusion approaches generally discard ancillary information produced by MVS algorithms (such as measurement confidence/uncertainty) that is otherwise extremely useful to enable adaptive fusion. To make use of such information, this paper proposes an efficient and scalable approach that incorporates the matching uncertainty to adaptively guide the fusion process. This seemingly straightforward idea has a higher-level advantage: first, the uncertainty information is obtained from global/semiglobal matching methods, which inherently populate global information of the scene, making the fusion process nonlocal. Secondly, these globally determined uncertainties are operated locally to achieve efficiency for processing large-sized images, making the method extremely practical to implement. The proposed method can exploit results from stereo pairs with small intersection angles to recover details for areas where dense buildings and narrow streets exist, but also to benefit from highly accurate 3D points generated in flat regions under large intersection angles. The proposed method was applied to DSMs generated from Worldview, GeoEye, and Pleiades stereo pairs covering a large area (400 km2). Experiments showed that we achieved an RMSE (root-mean-squared error) improvement of approximately 0.1–0.2 m over a typical Median Filter approach for fusion (equivalent to 5–10% of relative accuracy improvement).
Journal Article