Asset Details

MbrlCatalogueTitleDetail

Do you wish to reserve the book?

Speaker Localization Based on Audio-Visual Bimodal Fusion

by Jin, Hao-Ran , Zhu, Ying-Xin

in Accuracy / Algorithms / Audio equipment / Converters / Coordinates / Face recognition / Localization / Localization method / Mode localization / Redundancy

2021

Yes Please

Hey, we have placed the reservation for you!

By the way, why not check out events that you can attend while you pick your title.

Oops! Something went wrong.

Looks like we were not able to place the reservation. Kindly try again later.

Journal Article

Speaker Localization Based on Audio-Visual Bimodal Fusion

Jin, Hao-Ran,

Zhu, Ying-Xin

2021

Overview

The demand for fluency in human–computer interaction is on an increase globally; thus, the active localization of the speaker by the machine has become a problem worth exploring. Considering that the stability and accuracy of the single-mode localization method are low, while the multi-mode localization method can utilize the redundancy of information to improve accuracy and anti-interference, a speaker localization method based on voice and image multimodal fusion is proposed. First, the voice localization method based on time differences of arrival (TDOA) in a microphone array and the face detection method based on the AdaBoost algorithm are presented herein. Second, a multimodal fusion method based on spatiotemporal fusion of speech and image is proposed, and it uses a coordinate system converter and frame rate tracker. The proposed method was tested by positioning the speaker stand at 15 different points, and each point was tested 50 times. The experimental results demonstrate that there is a high accuracy when the speaker stands in front of the positioning system within a certain range.

Share this book

Add to My Shelf

Publisher

Fuji Technology Press Co. Ltd

Subject