O corpus AMR-PT e a anotação semântica de sentenças desafiadoras de textos jornalísticos e opinativos

Marcio Lima Inácio,Marco Antonio Sobrevilla Cabezudo,Ariani Di Felippo,Renata Ramisch,Thiago Alexandre Salgueiro Pardo

doi:10.1590/1678-460x202339355159

Abstract

ABSTRACT One of the most popular semantic representation languages in Natural Language Processing (NLP) is Abstract Meaning Representation (AMR). This formalism encodes the meaning of single sentences in directed rooted graphs. For English, there is a large annotated corpus that provides qualitative and reusable data for building or improving existing NLP methods and applications. For building AMR corpora for non-English languages, including Brazilian Portuguese, automatic and manual strategies have been conducted. The automatic annotation methods are essentially based on the cross-linguistic alignment of parallel corpora and the inheritance of the AMR annotation. The manual strategies focus on adapting the AMR English guidelines to a target language. Both annotation strategies have to deal with some phenomena that are challenging. This paper explores in detail some characteristics of Portuguese for which the AMR model had to be adapted and introduces two annotated corpora: AMRNews, a corpus of 870 annotated sentences from journalistic texts, and OpiSums-PT-AMR, comprising 404 opinionated sentences in AMR.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: DELTA: Documentação de Estudos em Lingüística Teórica e Aplicada	Publication Date: Jan 1, 2023
Citations: 1	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

O corpus AMR-PT e a anotação semântica de sentenças desafiadoras de textos jornalísticos e opinativos

Abstract

Talk to us

Similar Papers

More From: DELTA: Documentação de Estudos em Lingüística Teórica e Aplicada

Lead the way for us

Similar Papers

Domain-specific language models and lexicons for tagging
Anni R Coden ... Christopher G Chute
Journal of Biomedical Informatics | VOL. 38
Anni R Coden, et. al.Anni R Coden ... Christopher G Chute
02 Apr 2005
Journal of Biomedical Informatics | VOL. 38

Advanced Corpus Annotation Strategies for NLP. Applications in Automatic Summarization and Text Classiﬁcation

-

01 Jan 2020
01 Jan 2020

ScisummNet: A Large Annotated Corpus and Content-Impact Models for Scientific Paper Summarization with Citation Networks
Michihiro Yasunaga ... Irene Li
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 33
Michihiro Yasunaga, et. al.Michihiro Yasunaga ... Irene Li
17 Jul 2019
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 33

An Urdu semantic tagger - lexicons, corpora, methods and tools

-

30 Sep 2019
30 Sep 2019

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

O corpus AMR-PT e a anotação semântica de sentenças desafiadoras de textos jornalísticos e opinativos

Abstract

Talk to us

Similar Papers

More From: DELTA: Documentação de Estudos em Lingüística Teórica e Aplicada