Evaluation of Vector Transformations for Russian Word2Vec and FastText Embeddings

Olga Korogodina,Eduard Klyshinsky,Olesya Karpik

doi:10.51130/graphicon-2020-2-3-18

Abstract

Authors of Word2Vec claimed that their technology could solve the word analogy problem using the vector transformation in the introduced vector space. However, the practice demonstrates that it is not always true. In this paper, we investigate several Word2Vec and FastText model trained for the Russian language and find out reasons of such inconsistency. We found out that different types of words are demonstrating different behavior in the semantic space. FastText vectors are tending to find phonological analogies, while Word2Vec vectors are better in finding relations in geographical proper names. However, we found out that just four out of fifteen selected domains are demonstrating accuracy more that 0.8. We also draw a conclusion that in a common case, the task of word analogies could not be solved using a random word pair taken from two investigated categories. Our experiments have demonstrated that in some cases the length of the vectors could differ more than twice. Calculation of an average vector leads to a better solution here since it closer to more vectors.

Highlights

The basic point for the semantic space of natural language words was the paper [1] published in 2003
The FastText model was created to work with character n-grams and learn grammatical features of the language
We found several reasons why the vector transformation does not work on some categories of word analogies

Summary

Introduction

The basic point for the semantic space of natural language words was the paper [1] published in 2003 It introduced fixed-size vectors (embeddings) generated by a neural network using statistical information about the word context. This concept was developed in [2] where the author demonstrated that such pre-trained vectors can be useful for solution of different problems of natural language processing. The early experiments demonstrated that another favorite example, countries and their capitals, does not work correctly for any case – country, capital and pre-trained language model The accuracy of this analogy was pretty high but not enough to state that vector arithmetic works properly. The description of contextualized words embedding could be found in paper [9]

Formal Statement of the Problem of Word Analogies

Review of Affine Transformation Methods for the Problem of Word Analogies

Used Data Sets

Evaluation

Data Analysis

Discussion and Conclusion

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Proceedings of the 30th International Conference on Computer Graphics and Machine Vision (GraphiCon 2020). Part 2	Publication Date: Dec 17, 2020
Citations: 6	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Evaluation of Vector Transformations for Russian Word2Vec and FastText Embeddings

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Proceedings of the 30th International Conference on Computer Graphics and Machine Vision (GraphiCon 2020). Part 2

Lead the way for us

Similar Papers

A Well‐Rounded Diet: Fueling Children’s Vocabulary Development
Lori Bruner
The Reading Teacher | VOL. 74
Lori BrunerLori Bruner
27 Apr 2021
The Reading Teacher | VOL. 74

Modulating Effects of Contextual Emotions on the Neural Plasticity Induced by Word Learning.
Jingjing Guo ... Dingding Li
Frontiers in Human Neuroscience | VOL. 12
Jingjing Guo, et. al.Jingjing Guo ... Dingding Li
23 Nov 2018
Frontiers in Human Neuroscience | VOL. 12

The impact of lexical factors on children's word-finding errors.
Diane J German ... Rochelle S Newman
Journal of Speech, Language, and Hearing Research | VOL. 47
Diane J German, et. al.Diane J German ... Rochelle S Newman
01 Jun 2004
Journal of Speech, Language, and Hearing Research | VOL. 47

Slovene and Croatian word embeddings in terms of gender occupational analogies
Matej Ulčar ... Marko Robnik-Šikonja
Slovenščina 2.0: empirical, applied and interdisciplinary research | VOL. 9
Matej Ulčar, et. al.Matej Ulčar ... Marko Robnik-Šikonja
06 Jul 2021
Slovenščina 2.0: empirical, applied and interdisciplinary research | VOL. 9

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Evaluation of Vector Transformations for Russian Word2Vec and FastText Embeddings

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Proceedings of the 30th International Conference on Computer Graphics and Machine Vision (GraphiCon 2020). Part 2