Parsing Models for Identifying Multiword Expressions

Spence Green,Christopher D Manning,Marie-Catherine De Marneffe

doi:10.1162/coli_a_00139

Abstract

Multiword expressions lie at the syntax/semantics interface and have motivated alternative theories of syntax like Construction Grammar. Until now, however, syntactic analysis and multiword expression identification have been modeled separately in natural language processing. We develop two structured prediction models for joint parsing and multiword expression identification. The first is based on context-free grammars and the second uses tree substitution grammars, a formalism that can store larger syntactic fragments. Our experiments show that both models can identify multiword expressions with much higher accuracy than a state-of-the-art system based on word co-occurrence statistics. We experiment with Arabic and French, which both have pervasive multiword expressions. Relative to English, they also have richer morphology, which induces lexical sparsity in finite corpora. To combat this sparsity, we develop a simple factored lexical representation for the context-free parsing model. Morphological analyses are automatically transformed into rich feature tags that are scored jointly with lexical items. This technique, which we call a factored lexicon, improves both standard parsing and multiword expression identification accuracy.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Parsing Models for Identifying Multiword Expressions

Abstract

Talk to us

Similar Papers

More From: Computational Linguistics

Lead the way for us

Journal: Computational Linguistics	Publication Date: Mar 1, 2013
Citations: 83

Similar Papers

Putting the Horses Before the Cart: Identifying Multiword Expressions Before Translation
Carlos Ramisch
-
Carlos RamischCarlos Ramisch
01 Jan 2017
01 Jan 2017

Identification of Multiword Expressions in Technical Domains: Investigating Statistical and Alignment-Based Approaches
Aline Villavicencio ... Andre Machado
-
Aline Villavicencio, et. al.Aline Villavicencio ... Andre Machado
01 Sep 2009
01 Sep 2009

Identification of Nominal Multiword Expressions in Bengali using CRF
Tanmoy Chakraborty
-
Tanmoy ChakrabortyTanmoy Chakraborty
01 Dec 2012
01 Dec 2012

Semantic Extraction of Arabic Multiword Expressions
Samah Meghawry ... Tarek Elghazaly
-
Samah Meghawry, et. al.Samah Meghawry ... Tarek Elghazaly
23 Jan 2015
23 Jan 2015

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Parsing Models for Identifying Multiword Expressions

Abstract

Talk to us

Similar Papers

More From: Computational Linguistics