Asset Details
MbrlCatalogueTitleDetail
Do you wish to reserve the book?
Creation of annotated country-level dialectal Arabic resources: An unsupervised approach
by
Althobaiti, Maha J
in
Agreements
/ Algorithms
/ Arabic language
/ Corpus linguistics
/ Dialects
/ Identification
/ Language usage
/ Morphology
/ Natural language processing
/ Networking
/ Social networks
/ Speech
/ Standard dialects
/ Syntax
/ Words
/ Words (language)
2022
Hey, we have placed the reservation for you!
By the way, why not check out events that you can attend while you pick your title.
You are currently in the queue to collect this book. You will be notified once it is your turn to collect the book.
Oops! Something went wrong.
Looks like we were not able to place the reservation. Kindly try again later.
Are you sure you want to remove the book from the shelf?
Creation of annotated country-level dialectal Arabic resources: An unsupervised approach
by
Althobaiti, Maha J
in
Agreements
/ Algorithms
/ Arabic language
/ Corpus linguistics
/ Dialects
/ Identification
/ Language usage
/ Morphology
/ Natural language processing
/ Networking
/ Social networks
/ Speech
/ Standard dialects
/ Syntax
/ Words
/ Words (language)
2022
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
Do you wish to request the book?
Creation of annotated country-level dialectal Arabic resources: An unsupervised approach
by
Althobaiti, Maha J
in
Agreements
/ Algorithms
/ Arabic language
/ Corpus linguistics
/ Dialects
/ Identification
/ Language usage
/ Morphology
/ Natural language processing
/ Networking
/ Social networks
/ Speech
/ Standard dialects
/ Syntax
/ Words
/ Words (language)
2022
Please be aware that the book you have requested cannot be checked out. If you would like to checkout this book, you can reserve another copy
We have requested the book for you!
Your request is successful and it will be processed during the Library working hours. Please check the status of your request in My Requests.
Oops! Something went wrong.
Looks like we were not able to place your request. Kindly try again later.
Creation of annotated country-level dialectal Arabic resources: An unsupervised approach
Journal Article
Creation of annotated country-level dialectal Arabic resources: An unsupervised approach
2022
Request Book From Autostore
and Choose the Collection Method
Overview
The wide usage of multiple spoken Arabic dialects on social networking sites stimulates increasing interest in Natural Language Processing (NLP) for dialectal Arabic (DA). Arabic dialects represent true linguistic diversity and differ from modern standard Arabic (MSA). In fact, the complexity and variety of these dialects make it insufficient to build one NLP system that is suitable for all of them. In comparison with MSA, the available datasets for various dialects are generally limited in terms of size, genre and scope. In this article, we present a novel approach that automatically develops an annotated country-level dialectal Arabic corpus and builds lists of words that encompass 15 Arabic dialects. The algorithm uses an iterative procedure consisting of two main components: automatic creation of lists for dialectal words and automatic creation of annotated Arabic dialect identification corpus. To our knowledge, our study is the first of its kind to examine and analyse the poor performance of the MSA part-of-speech tagger on dialectal Arabic contents and to exploit that in order to extract the dialectal words. The pointwise mutual information association measure and the geographical frequency of word occurrence online are used to classify dialectal words. The annotated dialectal Arabic corpus (Twt15DA), built using our algorithm, is collected from Twitter and consists of 311,785 tweets containing 3,858,459 words in total. We randomly selected a sample of 75 tweets per country, 1125 tweets in total, and conducted a manual dialect identification task by native speakers. The results show an average inter-annotator agreement score equal to 64%, which reflects satisfactory agreement considering the overlapping features of the 15 Arabic dialects.
Publisher
Cambridge University Press
Subject
This website uses cookies to ensure you get the best experience on our website.