Asset Details

MbrlCatalogueTitleDetail

Do you wish to reserve the book?

Moving-Talker, Speaker-Independent Feature Study, and Baseline Results Using the CUAVE Multimodal Speech Corpus

by Gurbuz, Sabri , Gowdy, John N. , Tufekci, Zekeriya , Patterson, Eric K.

in Audio data / audio-visual speech recognition / Background noise / Formability / Image processing / multimodal database / Optical disks / Speech processing / speechreading / Strings

2002

Yes Please

Hey, we have placed the reservation for you!

By the way, why not check out events that you can attend while you pick your title.

Oops! Something went wrong.

Looks like we were not able to place the reservation. Kindly try again later.

Are you sure you want to remove the book from the shelf?

Moving-Talker, Speaker-Independent Feature Study, and Baseline Results Using the CUAVE Multimodal Speech Corpus

by Gurbuz, Sabri , Gowdy, John N. , Tufekci, Zekeriya , Patterson, Eric K.

in Audio data / audio-visual speech recognition / Background noise / Formability / Image processing / multimodal database / Optical disks / Speech processing / speechreading / Strings

2002

Confirm

Do you wish to request the book?

Moving-Talker, Speaker-Independent Feature Study, and Baseline Results Using the CUAVE Multimodal Speech Corpus

by Gurbuz, Sabri , Gowdy, John N. , Tufekci, Zekeriya , Patterson, Eric K.

in Audio data / audio-visual speech recognition / Background noise / Formability / Image processing / multimodal database / Optical disks / Speech processing / speechreading / Strings

2002

Please be aware that the book you have requested cannot be checked out. If you would like to checkout this book, you can reserve another copy

How would you like to get it?

Submit

We have requested the book for you!

Your request is successful and it will be processed during the Library working hours. Please check the status of your request in My Requests.

Oops! Something went wrong.

Looks like we were not able to place your request. Kindly try again later.

Journal Article

Moving-Talker, Speaker-Independent Feature Study, and Baseline Results Using the CUAVE Multimodal Speech Corpus

Gurbuz, Sabri,

Gowdy, John N.,

Tufekci, Zekeriya,

Patterson, Eric K.

2002

Overview

Strides in computer technology and the search for deeper, more powerful techniques in signal processing have brought multimodal research to the forefront in recent years. Audio-visual speech processing has become an important part of this research because it holds great potential for overcoming certain problems of traditional audio-only methods. Difficulties, due to background noise and multiple speakers in an application environment, are significantly reduced by the additional information provided by visual features. This paper presents information on a new audio-visual database, a feature study on moving speakers, and on baseline results for the whole speaker group. Although a few databases have been collected in this area, none has emerged as a standard for comparison. Also, efforts to date have often been limited, focusing on cropped video or stationary speakers. This paper seeks to introduce a challenging audio-visual database that is flexible and fairly comprehensive, yet easily available to researchers on one DVD. The Clemson University Audio-Visual Experiments (CUAVE) database is a speaker-independent corpus of both connected and continuous digit strings totaling over 7000 utterances. It contains a wide variety of speakers and is designed to meet several goals discussed in this paper. One of these goals is to allow testing of adverse conditions such as moving talkers and speaker pairs. A feature study of connected digit strings is also discussed. It compares stationary and moving talkers in a speaker-independent grouping. An image-processing-based contour technique, an image transform method, and a deformable template scheme are used in this comparison to obtain visual features. This paper also presents methods and results in an attempt to make these techniques more robust to speaker movement. Finally, initial baseline speaker-independent results are included using all speakers, and conclusions as well as suggested areas of research are given.

Share this book

Add to My Shelf

Publisher

Springer Nature B.V,SpringerOpen

Subject

Audio data

/ audio-visual speech recognition

/ Background noise

/ Formability

/ Image processing

/ multimodal database

/ Optical disks