Benchmarks for Pirá 2.0, a Reading Comprehension Dataset about the Ocean, the Brazilian Coast, and Climate Change

Paulo Pirozelli,Sarajane M Peres,Igor Silveira,Flávio Nakasato,Fabio G Cozman,Marcos M José,Anna H R Costa,Anarosa A F Brandão

doi:10.1162/dint_a_00245

Abstract

ABSTRACT Pirá is a reading comprehension dataset focused on the ocean, the Brazilian coast, and climate change, built from a collection of scientific abstracts and reports on these topics. This dataset represents a versatile language resource, particularly useful for testing the ability of current machine learning models to acquire expert scientific knowledge. Despite its potential, a detailed set of baselines has not yet been developed for Pirá. By creating these baselines, researchers can more easily utilize Pirá as a resource for testing machine learning models across a wide range of question answering tasks. In this paper, we define six benchmarks over the Pirá dataset, covering closed generative question answering, machine reading comprehension, information retrieval, open question answering, answer triggering, and multiple choice question answering. As part of this effort, we have also produced a curated version of the original dataset, where we fixed a number of grammar issues, repetitions, and other shortcomings. Furthermore, the dataset has been extended in several new directions, so as to face the aforementioned benchmarks: translation of supporting texts from English into Portuguese, classification labels for answerability, automatic paraphrases of questions and answers, and multiple choice candidates. The results described in this paper provide several points of reference for researchers interested in exploring the challenges provided by the Pirá dataset.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Benchmarks for Pirá 2.0, a Reading Comprehension Dataset about the Ocean, the Brazilian Coast, and Climate Change

Abstract

Talk to us

Similar Papers

More From: Data Intelligence

Lead the way for us

Journal: Data Intelligence	Publication Date: Mar 11, 2024
License type: CC BY 4.0

Similar Papers

MMM: Multi-Stage Multi-Task Learning for Multi-Choice Reading Comprehension
Di Jin ... Tagyoung Chung
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 34
Di Jin, et. al.Di Jin ... Tagyoung Chung
03 Apr 2020
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 34

Learning to Classify the Wrong Answers for Multiple Choice Question Answering (Student Abstract)
Hyeondey Kim ... Pascale Fung
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 34
Hyeondey Kim, et. al.Hyeondey Kim ... Pascale Fung
03 Apr 2020
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 34

Improving ranking-based question answering with weak supervision for low-resource Qur’anic texts
Mohammed Elkoumy ... Amany Sarhan
Artificial Intelligence Review | VOL. 58
Mohammed Elkoumy, et. al.Mohammed Elkoumy ... Amany Sarhan
14 Nov 2024
Artificial Intelligence Review | VOL. 58

Context Modeling with Evidence Filter for Multiple Choice Question Answering
Sicheng Yu ... Jing Jiang
-
Sicheng Yu, et. al.Sicheng Yu ... Jing Jiang
23 May 2022
23 May 2022

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Benchmarks for Pirá 2.0, a Reading Comprehension Dataset about the Ocean, the Brazilian Coast, and Climate Change

Abstract

Talk to us

Similar Papers

More From: Data Intelligence