Hierarchical Aggregation of Dialectal Data for Arabic Dialect Identification.

Nurpeiis Baimukan,Houda Bouamor,Nizar Habash

International Conference on Language Resources and Evaluation (LREC)（2022）

引用 0|浏览10

暂无评分

摘要

Arabic is a collection of dialectal variants that are historically related but significantly different. These differences can be seen across regions, countries, and even cities in the same countries. Previous work on Arabic Dialect identification has focused mainly on specific dialect levels (region, country, province, or city) using level-specific resources; and different efforts used different schemas and labels. In this paper, we present the first effort aiming at defining a standard unified three-level hierarchical schema (region-country-city) for dialectal Arabic classification. We map 29 different data sets to this unified schema, and use the common mapping to facilitate aggregating these data sets. We test the value of such aggregation by building language models and using them in dialect identification. We make our label mapping code and aggregated language models publicly available.

查看译文

关键词

Arabic Dialects, Dialect Identification, Language Models

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要