Asset Details

MbrlCatalogueTitleDetail

Do you wish to reserve the book?

The Development and Experimental Evaluation of a Multilingual Speech Corpus for Low-Resource Turkic Languages

by Tukeyev, Ualsher , Shormakova, Assem , Abduali, Balzhan , Amirova, Dina , Rakhimova, Diana , Aliyev, Rashid , Karibayeva, Aidana , Karyukin, Vladislav

in Accuracy / Artificial intelligence / Automatic text generation / Computational linguistics / Corpus linguistics / Datasets / Intonation / Kazakh language / Language / Language processing / Machine learning / Machine translation / Multilingualism / Natural language interfaces / Naturalness / parallel speech corpora / Phonetics / Recognition / speech corpus / Speech synthesis / Tatar / Text-to-speech / Transcription / Translating and interpreting / Turkic languages / Turkish / Turkish language / Uzbek / Uzbek language / Voice recognition / Word processing

2025

Yes Please

Hey, we have placed the reservation for you!

By the way, why not check out events that you can attend while you pick your title.

Oops! Something went wrong.

Looks like we were not able to place the reservation. Kindly try again later.

Are you sure you want to remove the book from the shelf?

The Development and Experimental Evaluation of a Multilingual Speech Corpus for Low-Resource Turkic Languages

by Tukeyev, Ualsher , Shormakova, Assem , Abduali, Balzhan , Amirova, Dina , Rakhimova, Diana , Aliyev, Rashid , Karibayeva, Aidana , Karyukin, Vladislav

2025

Confirm

Do you wish to request the book?

The Development and Experimental Evaluation of a Multilingual Speech Corpus for Low-Resource Turkic Languages

by Tukeyev, Ualsher , Shormakova, Assem , Abduali, Balzhan , Amirova, Dina , Rakhimova, Diana , Aliyev, Rashid , Karibayeva, Aidana , Karyukin, Vladislav

2025

Please be aware that the book you have requested cannot be checked out. If you would like to checkout this book, you can reserve another copy

How would you like to get it?

Submit

We have requested the book for you!

Your request is successful and it will be processed during the Library working hours. Please check the status of your request in My Requests.

Oops! Something went wrong.

Looks like we were not able to place your request. Kindly try again later.

Journal Article

The Development and Experimental Evaluation of a Multilingual Speech Corpus for Low-Resource Turkic Languages

Tukeyev, Ualsher,

Shormakova, Assem,

Abduali, Balzhan,

Amirova, Dina,

Rakhimova, Diana,

Aliyev, Rashid,

Karibayeva, Aidana,

Karyukin, Vladislav

2025

Overview

The development of parallel audio corpora for Turkic languages, such as Kazakh, Uzbek, and Tatar, remains a significant challenge in the development of multilingual speech synthesis, recognition systems, and machine translation. These languages are low-resource in speech technologies, lacking sufficiently large audio datasets with aligned transcriptions that are crucial for modern recognition, synthesis, and understanding systems. This article presents the development and experimental evaluation of a speech corpus focused on Turkic languages, intended for use in speech synthesis and automatic translation tasks. The primary objective is to create parallel audio corpora using a cascade generation method, which combines artificial intelligence and text-to-speech (TTS) technologies to generate both audio and text, and to evaluate the quality and suitability of the generated data. To evaluate the quality of synthesized speech, metrics measuring naturalness, intonation, expressiveness, and linguistic adequacy were applied. As a result, a multimodal (Kazakh–Turkish, Kazakh–Tatar, Kazakh–Uzbek) corpus was created, combining high-quality natural Kazakh audio with transcription and translation, along with synthetic audio in Turkish, Tatar, and Uzbek. These corpora offer a unique resource for speech and text processing research, enabling the integration of ASR, MT, TTS, and speech-to-speech translation (STS).

Share this book

Add to My Shelf

Publisher

MDPI AG

Subject

Accuracy

/ Artificial intelligence

/ Automatic text generation

/ Computational linguistics

/ Corpus linguistics

/ Datasets

/ Intonation