Massive vs. Curated Word Embeddings for Low-Resourced Languages. The Case of Yor\`ub\'a and Twi

Alabi, Jesujoba O.; Amponsah-Kaakyire, Kwabena; Adelani, David I.; España-Bonet, Cristina

Computer Science > Computation and Language

arXiv:1912.02481 (cs)

[Submitted on 5 Dec 2019 (v1), last revised 28 Mar 2020 (this version, v2)]

Title:Massive vs. Curated Word Embeddings for Low-Resourced Languages. The Case of Yorùbá and Twi

Authors:Jesujoba O. Alabi, Kwabena Amponsah-Kaakyire, David I. Adelani, Cristina España-Bonet

View PDF

Abstract:The success of several architectures to learn semantic representations from unannotated text and the availability of these kind of texts in online multilingual resources such as Wikipedia has facilitated the massive and automatic creation of resources for multiple languages. The evaluation of such resources is usually done for the high-resourced languages, where one has a smorgasbord of tasks and test sets to evaluate on. For low-resourced languages, the evaluation is more difficult and normally ignored, with the hope that the impressive capability of deep learning architectures to learn (multilingual) representations in the high-resourced setting holds in the low-resourced setting too. In this paper we focus on two African languages, Yorùbá and Twi, and compare the word embeddings obtained in this way, with word embeddings obtained from curated corpora and a language-dependent processing. We analyse the noise in the publicly available corpora, collect high quality and noisy data for the two languages and quantify the improvements that depend not only on the amount of data but on the quality too. We also use different architectures that learn word representations both from surface forms and characters to further exploit all the available information which showed to be important for these languages. For the evaluation, we manually translate the wordsim-353 word pairs dataset from English into Yorùbá and Twi. As output of the work, we provide corpora, embeddings and the test suits for both languages.

Comments:	9 pages, 4 tables. Accepted at LREC 2020
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:1912.02481 [cs.CL]
	(or arXiv:1912.02481v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1912.02481

Submission history

From: Cristina España-Bonet [view email]
[v1] Thu, 5 Dec 2019 10:25:32 UTC (28 KB)
[v2] Sat, 28 Mar 2020 21:38:21 UTC (29 KB)

Computer Science > Computation and Language

Title:Massive vs. Curated Word Embeddings for Low-Resourced Languages. The Case of Yorùbá and Twi

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Massive vs. Curated Word Embeddings for Low-Resourced Languages. The Case of Yorùbá and Twi

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators