Subwords-Only Alternatives to fastText for Morphologically Rich Languages

Tsolak Ghukasyan,Karen Avetisyan,Yeva Yeshilbashyan

doi:10.1134/s0361768821010059

Subwords-Only Alternatives to fastText for Morphologically Rich Languages

Tsolak Ghukasyan, Karen Avetisyan + Show 1 more

https://doi.org/10.1134/s0361768821010059

Copy DOI

Journal: Programming and Computer Software	Publication Date: Jan 1, 2021
Citations: 3

Affiliation: Russian-Armenian University

#Word-level Vectors #Subword Information + Show 8 more

Abstract
Full-Text PDF
Similar Papers

Abstract

In this work, we present purely subword-based alternatives to fastText word embedding algorithm The alternatives are modifications of the original fastText model, but rely on subword information only, eliminating the reliance on word-level vectors and at the same time helping to dramatically reduce the size of embeddings. Proposed models differ in their subword information extraction method: character n-grams, suffixes, and the byte-pair encoding units. We test the models in the task of morphological analysis and lemmatization for 3 morphologically rich languages: Finnish, Russian, and German. The results are compared with other recent subword-based models, demonstrating consistently higher results.

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Similar Papers

Paper Title

Journal

Date

Author

View more papers

More From: Programming and Computer Software

Paper Title

Journal

Date

Author

View more papers

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.