Asset Details
MbrlCatalogueTitleDetail
Do you wish to reserve the book?
Leveraging Multimodal Large Language Models (MLLMs) for Enhanced Object Detection and Scene Understanding in Thermal Images for Autonomous Driving Systems
by
Alhadidi, Taqwa I.
, Ashqar, Huthaifa I.
, Elhenawy, Mohammed
, Khanfar, Nour O.
in
Accuracy
/ Algorithms
/ Automation
/ autonomous driving systems
/ Autonomous vehicles
/ Cameras
/ Classification
/ Color imagery
/ Deep learning
/ Heat detection
/ Infrared imagery
/ Intelligent transportation systems
/ Language
/ Large language models
/ Machine learning
/ multimodal large language models (MLLMs)
/ Natural language
/ object detection
/ Object recognition
/ Recall
/ RGB
/ Scene analysis
/ Sensors
/ thermal images
/ Thermal imaging
2024
Hey, we have placed the reservation for you!
By the way, why not check out events that you can attend while you pick your title.
You are currently in the queue to collect this book. You will be notified once it is your turn to collect the book.
Oops! Something went wrong.
Looks like we were not able to place the reservation. Kindly try again later.
Are you sure you want to remove the book from the shelf?
Leveraging Multimodal Large Language Models (MLLMs) for Enhanced Object Detection and Scene Understanding in Thermal Images for Autonomous Driving Systems
by
Alhadidi, Taqwa I.
, Ashqar, Huthaifa I.
, Elhenawy, Mohammed
, Khanfar, Nour O.
in
Accuracy
/ Algorithms
/ Automation
/ autonomous driving systems
/ Autonomous vehicles
/ Cameras
/ Classification
/ Color imagery
/ Deep learning
/ Heat detection
/ Infrared imagery
/ Intelligent transportation systems
/ Language
/ Large language models
/ Machine learning
/ multimodal large language models (MLLMs)
/ Natural language
/ object detection
/ Object recognition
/ Recall
/ RGB
/ Scene analysis
/ Sensors
/ thermal images
/ Thermal imaging
2024
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
Do you wish to request the book?
Leveraging Multimodal Large Language Models (MLLMs) for Enhanced Object Detection and Scene Understanding in Thermal Images for Autonomous Driving Systems
by
Alhadidi, Taqwa I.
, Ashqar, Huthaifa I.
, Elhenawy, Mohammed
, Khanfar, Nour O.
in
Accuracy
/ Algorithms
/ Automation
/ autonomous driving systems
/ Autonomous vehicles
/ Cameras
/ Classification
/ Color imagery
/ Deep learning
/ Heat detection
/ Infrared imagery
/ Intelligent transportation systems
/ Language
/ Large language models
/ Machine learning
/ multimodal large language models (MLLMs)
/ Natural language
/ object detection
/ Object recognition
/ Recall
/ RGB
/ Scene analysis
/ Sensors
/ thermal images
/ Thermal imaging
2024
Please be aware that the book you have requested cannot be checked out. If you would like to checkout this book, you can reserve another copy
We have requested the book for you!
Your request is successful and it will be processed during the Library working hours. Please check the status of your request in My Requests.
Oops! Something went wrong.
Looks like we were not able to place your request. Kindly try again later.
Leveraging Multimodal Large Language Models (MLLMs) for Enhanced Object Detection and Scene Understanding in Thermal Images for Autonomous Driving Systems
Journal Article
Leveraging Multimodal Large Language Models (MLLMs) for Enhanced Object Detection and Scene Understanding in Thermal Images for Autonomous Driving Systems
2024
Request Book From Autostore
and Choose the Collection Method
Overview
The integration of thermal imaging data with multimodal large language models (MLLMs) offers promising advancements for enhancing the safety and functionality of autonomous driving systems (ADS) and intelligent transportation systems (ITS). This study investigates the potential of MLLMs, specifically GPT-4 Vision Preview and Gemini 1.0 Pro Vision, for interpreting thermal images for applications in ADS and ITS. Two primary research questions are addressed: the capacity of these models to detect and enumerate objects within thermal images, and to determine whether pairs of image sources represent the same scene. Furthermore, we propose a framework for object detection and classification by integrating infrared (IR) and RGB images of the same scene without requiring localization data. This framework is particularly valuable for enhancing the detection and classification accuracy in environments where both IR and RGB cameras are essential. By employing zero-shot in-context learning for object detection and the chain-of-thought technique for scene discernment, this study demonstrates that MLLMs can recognize objects such as vehicles and individuals with promising results, even in the challenging domain of thermal imaging. The results indicate a high true positive rate for larger objects and moderate success in scene discernment, with a recall of 0.91 and a precision of 0.79 for similar scenes. The integration of IR and RGB images further enhances detection capabilities, achieving an average precision of 0.93 and an average recall of 0.56. This approach leverages the complementary strengths of each modality to compensate for individual limitations. This study highlights the potential of combining advanced AI methodologies with thermal imaging to enhance the accuracy and reliability of ADS, while identifying areas for improvement in model performance.
Publisher
MDPI AG
Subject
This website uses cookies to ensure you get the best experience on our website.