Extraction of Relation Descriptors for Portuguese Using Conditional Random Fields

Sandra Collovini,Lucas Pugens,Renata Vieira,Aline A Vanin

doi:10.1007/978-3-319-12027-0_9

Abstract

AbstractAn important task in Information Extraction is Relation Extraction. Relation Extraction (RE) is the task of detecting and characterizing the semantic relations between entities in the text. This work proposes a new process for the extraction of any relation descriptors between Named Entities (NEs) in the Organization domain, for the Portuguese language, using the Conditional Random Fields (CRF) model. For example, from the following sentence fragment “Microsoft headquartered in Redmond, […]”, we can extract the relation descriptor “headquartered-in”, that relates the NEs “Microsoft” and “Redmond”. We evaluated different features configurations for CRF; the best results were obtained with the inclusion of the semantic feature based on the NE category, since this feature could express, in a better way, the kind of relationship between the pair of NEs we want to identify. The proposed process achieved F-measure rates of 45 % and 53 %, considering the extraction of complete and partial matching, respectively.KeywordsInformation extractionRelation extractionNamed entityNamed entity recognitionNatural language processingConditional random fields

Full Text