Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
14
result(s) for
"Koppal, Sanjeev J."
Sort by:
RD-GuideNet: A Depth-Guided Framework for Robust Detection, Segmentation, and Temporal Tracking of White Button Mushrooms
2026
Mushroom farms in the United States continue to face persistent labor shortages, especially during the harvesting of white button mushrooms (Agaricus bisporus) which requires selective picking by skilled workers. This study addresses this challenge by developing a depth-guided computer vision framework for automated mushroom detection, segmentation, and tracking to support timely harvest decisions, providing the foundation needed to support selective and timely robotic harvesting. The specific objectives of the study were to (1) develop a novel image-processing algorithm (RD-GuideNet) that integrates RGB and depth images for accurate detection and segmentation of mushrooms; (2) implement a custom depth-guided tracking algorithm to preserve mushroom identities across sequential frames; (3) compare the performance of RD-GuideNet against state-of-the-art deep learning models, YOLOv8 and YOLOv11, focusing on segmentation and tracking accuracies. The proposed RD-GuideNet achieved an F1-score of 0.93 for segmentation, outperforming YOLOv8 (0.88) and YOLOv11 (0.86), and produced sharper, more geometrically consistent boundaries that closely followed true mushroom cap contours. Its tracking consistency reached 92.7%, compared to YOLOv8 (95.3%) and YOLOv11 (94.6%). Although slightly lower, RD-GuideNet maintained high temporal consistency across dense mushroom beds. These results suggest that depth-based geometric reasoning and deep learning approaches exhibit complementary strengths in dense production scenes. Combining the two may further enhance detection reliability and shape fidelity, supporting high-precision perception for autonomous mushroom harvesting. A comprehensive quantitative evaluation of such a hybrid framework will be investigated in future work.
Journal Article
Exploiting DLP Illumination Dithering for Reconstruction and Photography of High-Speed Scenes
by
Narasimhan, Srinivasa G.
,
Koppal, Sanjeev J.
,
Yamazaki, Shuntaro
in
Analysis
,
Applied sciences
,
Artificial Intelligence
2012
In this work, we recover fast moving scenes by exploiting the high-speed illumination “dithering” of cheap and easily available digital light processing (DLP) projectors. We first show how to reverse-engineer the temporal dithering for off-the-shelf projectors, using a high-speed camera. DLP dithering can produce temporal patterns commonly used in active vision techniques. Since the dithering occurs at a very high frame-rate, such illumination-based methods can be “speed up” for fast scenes. We demonstrate this with three applications, each of which only requires a single slide to be displayed by the DLP projector. The quality of the result is determined by the camera frame-rate available to the user. Pairing a high-speed camera and a DLP projector, we demonstrate structured light reconstruction at 100 Hz. With the same camera and three or more DLP projectors, we show photometric stereo and demultiplexing applications at 300 Hz. Finally, with a real-time (60 Hz) or still camera, we show that DLP illumination acts as a very fast flash, allowing strobe photography of high-speed scenes. We discuss, in depth, some characteristics of the temporal dithering with a case study of a particular projector. Finally, we describe limitations, trade-offs and other issues relating to this work.
Journal Article
Schrödinger's Camera: First Steps Towards a Quantum-Based Privacy Preserving Camera
2023
Privacy-preserving vision must overcome the dual challenge of utility and privacy. Too much anonymity renders the images useless, but too little privacy does not protect sensitive data. We propose a novel design for privacy preservation, where the imagery is stored in quantum states. In the future, this will be enabled by quantum imaging cameras, and, currently, storing very low resolution imagery in quantum states is possible. Quantum state imagery has the advantage of being both private and non-private till the point of measurement. This occurs even when images are manipulated, since every quantum action is fully reversible. We propose a control algorithm, based on double deep Q-learning, to learn how to anonymize the image before measurement. After learning, the RL weights are fixed, and new attack neural networks are trained from scratch to break the system's privacy. Although all our results are in simulation, we demonstrate, with these first steps, that it is possible to control both privacy and utility in a quantum-based manner.
FoveaSPAD: Exploiting Depth Priors for Adaptive and Efficient Single-Photon 3D Imaging
2024
Fast, efficient, and accurate depth-sensing is important for safety-critical applications such as autonomous vehicles. Direct time-of-flight LiDAR has the potential to fulfill these demands, thanks to its ability to provide high-precision depth measurements at long standoff distances. While conventional LiDAR relies on avalanche photodiodes (APDs), single-photon avalanche diodes (SPADs) are an emerging image-sensing technology that offer many advantages such as extreme sensitivity and time resolution. In this paper, we remove the key challenges to widespread adoption of SPAD-based LiDARs: their susceptibility to ambient light and the large amount of raw photon data that must be processed to obtain in-pixel depth estimates. We propose new algorithms and sensing policies that improve signal-to-noise ratio (SNR) and increase computing and memory efficiency for SPAD-based LiDARs. During capture, we use external signals to foveate, i.e., guide how the SPAD system estimates scene depths. This foveated approach allows our method to ``zoom into'' the signal of interest, reducing the amount of raw photon data that needs to be stored and transferred from the SPAD sensor, while also improving resilience to ambient light. We show results both in simulation and also with real hardware emulation, with specific implementations achieving a 1548-fold reduction in memory usage, and our algorithms can be applied to newly available and future SPAD arrays.
Design of an Adaptive Lightweight LiDAR to Decouple Robot-Camera Geometry
2024
A fundamental challenge in robot perception is the coupling of the sensor pose and robot pose. This has led to research in active vision where robot pose is changed to reorient the sensor to areas of interest for perception. Further, egomotion such as jitter, and external effects such as wind and others affect perception requiring additional effort in software such as image stabilization. This effect is particularly pronounced in micro-air vehicles and micro-robots who typically are lighter and subject to larger jitter but do not have the computational capability to perform stabilization in real-time. We present a novel microelectromechanical (MEMS) mirror LiDAR system to change the field of view of the LiDAR independent of the robot motion. Our design has the potential for use on small, low-power systems where the expensive components of the LiDAR can be placed external to the small robot. We show the utility of our approach in simulation and on prototype hardware mounted on a UAV. We believe that this LiDAR and its compact movable scanning design provide mechanisms to decouple robot and sensor geometry allowing us to simplify robot perception. We also demonstrate examples of motion compensation using IMU and external odometry feedback in hardware.
VL-Explore: Zero-shot Vision-Language Exploration and Target Discovery by Mobile Robots
by
Islam, Md Jahidul
,
Abdullah, Adnan
,
Zhang, Yuxuan
in
Algorithms
,
Monocular vision
,
Navigation
2025
Vision-language navigation (VLN) has emerged as a promising paradigm, enabling mobile robots to perform zero-shot inference and execute tasks without specific pre-programming. However, current systems often separate map exploration and path planning, with exploration relying on inefficient algorithms due to limited (partially observed) environmental information. In this paper, we present a novel navigation pipeline named \"VL-Explore\" for simultaneous exploration and target discovery in unknown environments, leveraging the capabilities of a vision-language model named CLIP. Our approach requires only monocular vision and operates without any prior map or knowledge about the target. For comprehensive evaluations, we designed a functional prototype of a UGV (unmanned ground vehicle) system named \"Open Rover\", a customized platform for general-purpose VLN tasks. We integrated and deployed the VL-Explore pipeline on Open Rover to evaluate its throughput, obstacle avoidance capability, and trajectory performance across various real-world scenarios. Experimental results demonstrate that VL-Explore consistently outperforms traditional map-traversal algorithms and achieves performance comparable to path-planning methods that depend on prior map and target knowledge. Notably, VL-Explore offers real-time active navigation without requiring pre-captured candidate images or pre-built node graphs, addressing key limitations of existing VLN pipelines.
SaccadeCam: Adaptive Visual Attention for Monocular Depth Sensing
2021
Most monocular depth sensing methods use conventionally captured images that are created without considering scene content. In contrast, animal eyes have fast mechanical motions, called saccades, that control how the scene is imaged by the fovea, where resolution is highest. In this paper, we present the SaccadeCam framework for adaptively distributing resolution onto regions of interest in the scene. Our algorithm for adaptive resolution is a self-supervised network and we demonstrate results for end-to-end learning for monocular depth estimation. We also show preliminary results with a real SaccadeCam hardware prototype.
Robotic Monitoring of Colorimetric Leaf Sensors for Precision Agriculture
2025
Common remote sensing modalities (RGB, multispectral, hyperspectral imaging or LiDAR) are often used to indirectly measure crop health and do not directly capture plant stress indicators. Commercially available direct leaf sensors are bulky, powered electronics that are expensive and interfere with crop growth. In contrast, low-cost, passive and bio-degradable leaf sensors offer an opportunity to advance real-time monitoring as they directly interface with the crop surface while not interfering with crop growth. To this end, we co-design a sensor-detector system, where the sensor is a passive colorimetric leaf sensor that directly measures crop health in a precision agriculture setting, and the detector autonomously obtains optical signals from these leaf sensors. The detector comprises a low size weight and power (SWaP) mobile ground robot with an onboard monocular RGB camera and object detector to localize each leaf sensor, as well as a hyperspectral camera with a motorized mirror and halogen light to acquire hyperspectral images. The sensor's crop health-dependent optical signals can be extracted from the hyperspectral images. The proof-of-concept system is demonstrated in row-crop environments both indoors and outdoors where it is able to autonomously navigate, locate and obtain a hyperspectral image of all leaf sensors present, and acquire interpretable spectral resonance with 80 \\(\\%\\) accuracy within a required retrieval distance from the sensor.
Learning Privacy Preserving Encodings through Adversarial Training
by
Pittaluga, Francesco
,
Koppal, Sanjeev J
,
Chakrabarti, Ayan
in
Coding
,
Inference
,
Neural networks
2018
We present a framework to learn privacy-preserving encodings of images (or other high-dimensional data) to inhibit inference of a chosen private attribute. Rather than encoding a fixed dataset or inhibiting a fixed estimator, we aim to to learn an encoding function such that even after this function is fixed, an estimator with knowledge of the encoding is unable to learn to accurately predict the private attribute, when generalizing beyond a training set. We formulate this as adversarial optimization of an encoding function against a classifier for the private attribute, with both modeled as deep neural networks. We describe an optimization approach which successfully yields an encoder that permanently limits inference of the private attribute, while preserving either a generic notion of information, or the estimation of a different, desired, attribute. We experimentally validate the efficacy of our approach on private tasks of real-world complexity, by learning to prevent detection of scene classes from the Places-365 dataset.
HyperPRI: A Dataset of Hyperspectral Images for Underground Plant Root Study
2024
Collecting and analyzing hyperspectral imagery (HSI) of plant roots over time can enhance our understanding of their function, responses to environmental factors, turnover, and relationship with the rhizosphere. Current belowground red-green-blue (RGB) root imaging studies infer such functions from physical properties like root length, volume, and surface area. HSI provides a more complete spectral perspective of plants by capturing a high-resolution spectral signature of plant parts, which have extended studies beyond physical properties to include physiological properties, chemical composition, and phytopathology. Understanding crop plants’ physical, physiological, and chemical properties enables researchers to determine high-yielding, drought-resilient genotypes that can withstand climate changes and sustain future population needs. However, most HSI plant studies use cameras positioned above ground, and thus, similar belowground advances are urgently needed. One reason for the sparsity of belowground HSI studies is that root features often have limited distinguishing reflectance intensities compared to surrounding soil, potentially rendering conventional image analysis methods ineffective. Here we present HyperPRI, a novel dataset containing RGB and HSI data for in situ, non-destructive, underground plant root analysis using ML tools. HyperPRI contains images of plant roots grown in rhizoboxes for two annual crop species – peanut (Arachis hypogaea) and sweet corn (Zea mays). Drought conditions are simulated once, and the boxes are imaged and weighed on select days across two months. Along with the images, we provide hand-labeled semantic masks and imaging environment metadata. Additionally, we present baselines for root segmentation on this dataset and draw comparisons between methods that focus on spatial, spectral, and spatialspectral features to predict the pixel-wise labels. Results demonstrate that combining HyperPRI’s hyperspectral and spatial information improves semantic segmentation of target objects.