Asset Details
MbrlCatalogueTitleDetail
Do you wish to reserve the book?
Speaker Localization Based on Audio-Visual Bimodal Fusion
by
Jin, Hao-Ran
, Zhu, Ying-Xin
in
Accuracy
/ Algorithms
/ Audio equipment
/ Converters
/ Coordinates
/ Face recognition
/ Localization
/ Localization method
/ Mode localization
/ Redundancy
2021
Hey, we have placed the reservation for you!
By the way, why not check out events that you can attend while you pick your title.
You are currently in the queue to collect this book. You will be notified once it is your turn to collect the book.
Oops! Something went wrong.
Looks like we were not able to place the reservation. Kindly try again later.
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
Do you wish to request the book?
Speaker Localization Based on Audio-Visual Bimodal Fusion
by
Jin, Hao-Ran
, Zhu, Ying-Xin
in
Accuracy
/ Algorithms
/ Audio equipment
/ Converters
/ Coordinates
/ Face recognition
/ Localization
/ Localization method
/ Mode localization
/ Redundancy
2021
Please be aware that the book you have requested cannot be checked out. If you would like to checkout this book, you can reserve another copy
We have requested the book for you!
Your request is successful and it will be processed during the Library working hours. Please check the status of your request in My Requests.
Oops! Something went wrong.
Looks like we were not able to place your request. Kindly try again later.
Journal Article
Speaker Localization Based on Audio-Visual Bimodal Fusion
2021
Request Book From Autostore
and Choose the Collection Method
Overview
The demand for fluency in human–computer interaction is on an increase globally; thus, the active localization of the speaker by the machine has become a problem worth exploring. Considering that the stability and accuracy of the single-mode localization method are low, while the multi-mode localization method can utilize the redundancy of information to improve accuracy and anti-interference, a speaker localization method based on voice and image multimodal fusion is proposed. First, the voice localization method based on time differences of arrival (TDOA) in a microphone array and the face detection method based on the AdaBoost algorithm are presented herein. Second, a multimodal fusion method based on spatiotemporal fusion of speech and image is proposed, and it uses a coordinate system converter and frame rate tracker. The proposed method was tested by positioning the speaker stand at 15 different points, and each point was tested 50 times. The experimental results demonstrate that there is a high accuracy when the speaker stands in front of the positioning system within a certain range.
Publisher
Fuji Technology Press Co. Ltd
Subject
This website uses cookies to ensure you get the best experience on our website.