Unsupervised models for morpheme segmentation and morphology learning

Mathias Creutz,Krista Lagus

doi:10.1145/1187415.1187418

Unsupervised models for morpheme segmentation and morphology learning

Mathias Creutz, Krista Lagus

https://doi.org/10.1145/1187415.1187418

Copy DOI

Journal: ACM Transactions on Speech and Language Processing	Publication Date: Jan 1, 2007
Citations: 365

Affiliation: University of Helsinki

#Raw Text Data #Morpheme Segmentation + Show 8 more

Abstract
Full-Text PDF
Similar Papers

Abstract

We present a model family called Morfessor for the unsupervised induction of a simple morphology from raw text data. The model is formulated in a probabilistic maximum a posteriori framework. Morfessor can handle highly inflecting and compounding languages where words can consist of lengthy sequences of morphemes. A lexicon of word segments, called morphs , is induced from the data. The lexicon stores information about both the usage and form of the morphs. Several instances of the model are evaluated quantitatively in a morpheme segmentation task on different sized sets of Finnish as well as English data. Morfessor is shown to perform very well compared to a widely known benchmark algorithm, in particular on Finnish data.

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Similar Papers

Paper Title

Journal

Date

Author

View more papers

More From: ACM Transactions on Speech and Language Processing

Paper Title

Journal

Date

Author

View more papers

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.