REAL

Investigating Sub-Word Embedding Strategies for the Morphologically Rich and Free Phrase-Order Hungarian

Döbrössy, Bálint and Makrai, Márton and Tarján, Balázs and Szaszák, György (2019) Investigating Sub-Word Embedding Strategies for the Morphologically Rich and Free Phrase-Order Hungarian. In: 4th Workshop on Representation Learning for NLP (RepL4NLP-2019), 2019.08.02, Florence, Italy.

[img]
Preview
Text
W194321.pdf

Download (253kB) | Preview

Abstract

For morphologically rich languages, word embeddings provide less consistent semantic representations due to higher variance in word forms. Moreover, these languages often allow for less constrained word order, which further increases variance. For the highly agglutinative Hungarian, semantic accuracy of word embeddings measured on word analogy tasks drops by 50-75% compared to English. We observed that embeddings learn morphosyntax quite well instead. Therefore, we explore and evaluate several sub-word unit based embedding strategies – character n-grams, lemmatization provided by an NLP-pipeline, and segments obtained in unsupervised learning (morfessor) – to boost semantic consistency in Hungarian word vectors. The effect of changing embedding dimension and context window size have also been considered. Morphological analysis based lemmatization was found to be the best strategy to improve embeddings’ semantic accuracy, whereas adding character n-grams was found consistently counterproductive in this regard.

Item Type: Conference or Workshop Item (Paper)
Subjects: P Language and Literature / nyelvészet és irodalom > P0 Philology. Linguistics / filológia, nyelvészet
SWORD Depositor: MTMT SWORD
Depositing User: MTMT SWORD
Date Deposited: 27 Sep 2019 09:04
Last Modified: 17 Apr 2023 14:13
URI: http://real.mtak.hu/id/eprint/101686

Actions (login required)

Edit Item Edit Item