Evaluating Surprise Adequacy for Question Answering

Seah Kim,Shin Yoo

doi:10.1145/3387940.3391465

Evaluating Surprise Adequacy for Question Answering

Seah Kim, Shin Yoo

https://doi.org/10.1145/3387940.3391465

Copy DOI

Publication Date: Jun 27, 2020

Citations: 10

Affiliation: Korea Advanced Institute of Science and Technology

#Adequacy Metrics #Stanford Question Answering Dataset + Show 8 more

Abstract
Full-Text
Similar Papers

Abstract

With the wide and rapid adoption of Deep Neural Networks (DNNs) in various domains, an urgent need to validate their behaviour has risen, resulting in various test adequacy metrics for DNNs. One of the metrics, Surprise Adequacy (SA), aims to measure how surprising a new input is based on the similarity to the data used for training. While SA has been evaluated to be effective for image classifiers based on Convolutional Neural Networks (CNNs), it has not been studied for the Natural Language Processing (NLP) domain. This paper applies SA to NLP, in particular to the question answering task: the aim is to investigate whether SA correlates well with the correctness of answers. An empirical evaluation using the widely used Stanford Question Answering Dataset (SQuAD) shows that SA can work well as a test adequacy metric for the question answering task.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Similar Papers

Paper Title

Journal

Date

Author

View more papers

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.