UnibucKernel Reloaded: First Place in Arabic Dialect Identification for the Second Year in a Row

Butnaru, Andrei M.; Ionescu, Radu Tudor

Computer Science > Computation and Language

arXiv:1805.04876 (cs)

[Submitted on 13 May 2018 (v1), last revised 28 Jul 2018 (this version, v4)]

Title:UnibucKernel Reloaded: First Place in Arabic Dialect Identification for the Second Year in a Row

Authors:Andrei M. Butnaru, Radu Tudor Ionescu

View PDF

Abstract:We present a machine learning approach that ranked on the first place in the Arabic Dialect Identification (ADI) Closed Shared Tasks of the 2018 VarDial Evaluation Campaign. The proposed approach combines several kernels using multiple kernel learning. While most of our kernels are based on character p-grams (also known as n-grams) extracted from speech or phonetic transcripts, we also use a kernel based on dialectal embeddings generated from audio recordings by the organizers. In the learning stage, we independently employ Kernel Discriminant Analysis (KDA) and Kernel Ridge Regression (KRR). Preliminary experiments indicate that KRR provides better classification results. Our approach is shallow and simple, but the empirical results obtained in the 2018 ADI Closed Shared Task prove that it achieves the best performance. Furthermore, our top macro-F1 score (58.92%) is significantly better than the second best score (57.59%) in the 2018 ADI Shared Task, according to the statistical significance test performed by the organizers. Nevertheless, we obtain even better post-competition results (a macro-F1 score of 62.28%) using the audio embeddings released by the organizers after the competition. With a very similar approach (that did not include phonetic features), we also ranked first in the ADI Closed Shared Tasks of the 2017 VarDial Evaluation Campaign, surpassing the second best method by 4.62%. We therefore conclude that our multiple kernel learning method is the best approach to date for Arabic dialect identification.

Comments:	This paper presents the UnibucKernel team's participation at the 2018 Arabic Dialect Identification Shared Task. Accepted at the VarDial Workshop of COLING 2018
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:1805.04876 [cs.CL]
	(or arXiv:1805.04876v4 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1805.04876

Submission history

From: Radu Tudor Ionescu [view email]
[v1] Sun, 13 May 2018 12:53:47 UTC (39 KB)
[v2] Fri, 25 May 2018 12:48:33 UTC (40 KB)
[v3] Thu, 21 Jun 2018 16:24:46 UTC (40 KB)
[v4] Sat, 28 Jul 2018 11:03:54 UTC (40 KB)

Computer Science > Computation and Language

Title:UnibucKernel Reloaded: First Place in Arabic Dialect Identification for the Second Year in a Row

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:UnibucKernel Reloaded: First Place in Arabic Dialect Identification for the Second Year in a Row

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators